76 lines
2.8 KiB
Markdown
Raw Normal View History

# Architecture
Nimcheck uses a four-layer pipeline architecture. Each layer is independent and communicates through well-defined interfaces.
```
Source Code (string or file)
|
v
+-------------------+
| Detector | LanguageFlavor from extension, shebang, or content analysis
+-------------------+
|
v
+-------------------+
| Tokenizer | Converts raw source into a seq[Token] stream
+-------------------+
|
v
+-------------------+
| Validator | Analyzes token stream for syntax errors and structure
+-------------------+
|
v
+-------------------+
| Reporting | Formats errors as human-readable text or structured JSON
+-------------------+
|
v
ValidationResult
```
### Layer 1: Detector (`src/nimcheck/core/detector.nim`)
Language auto-detection with cascading priority:
| Priority | Method | Source | Confidence |
|----------|--------|--------|------------|
| 1 (highest) | Explicit | User-provided `LanguageFlavor` | 1.0 |
| 2 | File extension | 70+ extension-to-flavor mappings | 0.9 |
| 3 | Shebang line | `#!` line interpreter matching (bash, python, node, ruby, lua, swift) | 0.95 |
| 4 | Content analysis | Keyword scoring, structural patterns, bracket usage | 0.6-0.9 |
The `detectFlavor()` function returns a `(LanguageFlavor, FlavorDetectionMethod, confidence: float)` tuple.
### Layer 2: Tokenizer (`src/nimcheck/core/tokenizerbase.nim`, per-language impls)
Each tokenizer converts raw source into a `seq[Token]`. Tokenizers track:
- **Position**: line, column, and byte offset for every token
- **Bracket stack**: `(`, `[`, `{`, `<` pushed and popped with position tracking
- **String context**: escape sequences, multi-line delimiters, interpolation
- **Comment context**: line comments, block comments with nesting support
Tokenizers emit errors for:
- Unclosed strings (`ErrUnclosedString`)
- Unclosed comments (`ErrUnclosedComment`)
- Unexpected characters (`ErrUnexpectedToken`)
- Unclosed block constructs (language-specific, e.g., `${}`, `$()`)
### Layer 3: Validator (`src/nimcheck/core/validatorbase.nim`, per-language impls)
Each validator:
1. Creates its tokenizer and runs `tokenize()`
2. Collects tokenizer errors
3. Runs `analyzeTokens()` - language-specific structural analysis
4. Runs `buildModuleInfo()` - populates `ModuleInfo` with imports, functions, classes
5. Computes validity, warning/info counts, duration
Validators register themselves via a factory pattern at import time. The registry maps flavor strings (e.g., `"python"`, `"bash"`) to constructors.
### Layer 4: Reporting (`src/nimcheck/reporting/errors.nim`)
Produces:
- **Human-readable reports**: `formatReport()` with position, severity, code, message, source context with carets, and hints