Detectors and patterns
What the engine and ruleset each own
| Engine | Ruleset |
|---|---|
| Tokenizer grammar, token fields, contexts and option validation. | Keyword lexicons and reusable value lists. |
| Bounded pattern compiler and matcher. | Token patterns, predicates and diagnostic reasons. |
| SQL and HTML fingerprint interpretation. | Fingerprint lexicons, fingerprint lists and reported reason. |
| Work budgets and version checks. | Rule targets, transformations, score and detector selection. |
The available grammar is listed in Tokenizers. A ruleset can add words and patterns without adding executable tokenizer code.
Lists and lexicons
A lists entry has an id and non-empty values. A ruleset can contain up to 64 lists, each with at most 16,384 values of at most 256 characters. Operators and pattern predicates refer to lists by ID within the same ruleset.
A lexicons entry has an id, tokenizer and entries, mapping a word to up to eight classes. There can be at most 16 lexicons and 16,384 entries per lexicon. In pattern mode, class names match [A-Z][A-Za-z0-9]* and cannot replace tokenizer grammar classes. SQL fingerprint lexicons instead use the permitted libinjection type letters. A lexicon cannot mix incompatible uses.
Detector definitions
A detector declares id, tokenizer, mode, optional lexicon, and optional options. Patterns mode requires a non-empty patterns array and forbids fingerprint fields. Fingerprint mode uses fingerprints to name a list and requires fingerprintReason; it has no patterns. The tokenizer determines whether fingerprint mode is available and which options and contexts are valid.
There can be up to 32 detectors, 256 patterns per detector, 16 steps per pattern and four captures per pattern. A fingerprint list is subject to the 16,384-value list bound. These are admission limits; runtime work budgets still apply.
A rule calls a same-ruleset detector with operator: { "kind": "Detect", "value": "detector-id" }. Detect accepts only its supported untrusted value targets, not arbitrary facts or metadata. It cannot be negated and cannot take values, number, list or an operand. Operators and Targets describe the available surface.
Pattern language
Each pattern has a diagnostic reason and a sequence. A step’s class array is a choice of grammar or lexicon classes. Without repeat, the step consumes one matching token; ? makes it optional and * permits zero through 32 repetitions. At least one step must require a token. capture names a token for later predicates.
skip permits listed token classes between steps. reset discards progress on listed classes. Without anchor, matching can begin later in the token stream; anchor: "Start" restricts the start. where on a step tests the token’s text or a field defined by its tokenizer. Supported checks are membership in a named list (in), literal prefix with optional minLength, a list-driven prefixAnyIgnoringWhitespace, startsWithAny, or nonBlank.
Pattern-level where compares captures with same, or tests a capture’s class using class. Captures and their references must be valid at compilation. See the SQL, XSS and shell examples for complete syntax.
Matching and cost
The compiled automaton processes tokens without regex-style backtracking. It reports a pattern completing at the earliest token; simultaneous completions use pattern list order. Detector contexts run in configured order until one matches. Tokenization cost is charged for each attempted context, and pattern work is charged as tokens multiplied by compiled positions.
Fingerprint mode implements the pinned libinjection v4.0.0-compatible SQL and XSS decisions. Compatibility describes that algorithm, not proof that every injection is detected. The third-party notices give its license, and protection coverage reports the measured coverage results.