October 10, 2026
Detection as Code Meets Synthetic Attack Logs: Turning Validation into a CI/CD Test Oracle
Part 2 of a 3-part series on modern detection engineering. Part 1 introduced a Detection-as-Code pipeline for Splunk correlation rulesβ¦

By Pietro Romano / SecBeret
5 min read
Part 2 of a 3-part series on modern detection engineering. Part 1 introduced a Detection-as-Code pipeline for Splunk correlation rules. This part plugs the validation layer into it, building on Synthetic Attack Log Generation for Splunk. Part 3 covers AI agents that author rules inside the same pipeline.
1. Introduction and Scope
The Part 1 pipeline has a weak stage. After a Sigma rule is linted and compiled to SPL, the "simulate" step assumes either a live Attack Range run or a real incident. Both are too slow and too costly to run on every pull request, and a test that is skipped under time pressure is not a control.
The synthetic generation framework already solves the data problem: it can produce attack telemetry for a given TTP without executing the attack. What it lacks, as a standalone methodology, is a contract that lets a CI job call it. This article defines that contract.
Scope:
- Splunk Enterprise / Cloud with ES, correlation rules compiled from Sigma
- Strategy A (replay) and Strategy B (LLM-generated events via HEC) as interchangeable data providers
- CI/CD integration, regression testing, and confidence scoring
Out of scope:
- Agent-driven rule authoring (Part 3)
- UEBA and baseline tuning
- Any execution of real attack techniques in production
2. From Framework to Test Oracle
In the original article, generation is something an engineer runs when they want to test a technique. In a DaC pipeline it has to behave like a unit test dependency: invoked automatically, deterministic in its verdict, and fast enough not to block the merge.
The shift is a change of interface, not of method. The oracle takes a rule and a set of techniques, and returns a verdict:
// request
{
"rule_id": "8f5c1a2e-7b3d-4e91-9c6a-2d4f8e1a9b7c",
"techniques": ["T1059.001"],
"provider": "replay",
"runs": 5
}
// response
{
"rule_id": "8f5c1a2e-7b3d-4e91-9c6a-2d4f8e1a9b7c",
"results": [
{"technique": "T1059.001", "runs": 5, "fired": 5, "verdict": "pass"}
],
"negative_results": {"runs": 3, "fired": 0, "verdict": "pass"}
}// request
{
"rule_id": "8f5c1a2e-7b3d-4e91-9c6a-2d4f8e1a9b7c",
"techniques": ["T1059.001"],
"provider": "replay",
"runs": 5
}
// response
{
"rule_id": "8f5c1a2e-7b3d-4e91-9c6a-2d4f8e1a9b7c",
"results": [
{"technique": "T1059.001", "runs": 5, "fired": 5, "verdict": "pass"}
],
"negative_results": {"runs": 3, "fired": 0, "verdict": "pass"}
}The negative_results block matters as much as the positive one. A rule that fires on everything passes every positive test.
3. Coverage-Driven Test Selection
Generating data for the whole ATT&CK catalog on every commit is wasteful. The Sigma file already declares what the rule claims to cover:
tags:
- attack.execution
- attack.t1059.001tags:
- attack.execution
- attack.t1059.001The CI job extracts the attack.tNNNN tags from the changed files and passes only those techniques to the oracle. The rule's declared coverage becomes its test plan. Two consequences follow:
- A rule with no technique tag cannot be tested, so the metadata check from Part 1 fails the PR before the oracle is called.
- A rule that claims a technique the oracle has no dataset for fails explicitly, instead of passing by omission. That failure is a coverage gap report in disguise.
4. The Regression Suite
Testing the changed rule is necessary and insufficient. Many failures are indirect: a change to a field mapping or an eventtype alters what a different rule sees.
The approach is to keep a persistent library of datasets, one per covered technique, and re-run it against every rule that shares an affected CIM data model. If a PR touches something feeding Endpoint.Processes, the regression run covers all rules built on that data model, not only the one in the diff.
Maintaining this library is a cost, and it is worth saying so plainly. Strategy A datasets are stable but age as attacker tradecraft changes. Strategy B datasets can be regenerated, but a regression suite whose inputs drift between runs is no longer a regression suite. The practical compromise is to freeze generated events once validated, store them in the repository next to the rule, and treat regeneration as a deliberate, reviewed change.
5. Beyond Pass/Fail: Confidence Scoring
Replay is deterministic. LLM generation is not: two calls for the same TTP can produce events that differ in command-line arguments, parent processes, or field completeness. A single pass on a single generation says little.
Running N generations per technique and recording the pass rate gives a usable signal:
def confidence(fired: int, runs: int) -> str:
rate = fired / runs
if runs < 5:
return "insufficient-runs"
if rate == 1.0:
return "high"
if rate >= 0.8:
return "medium" # review: rule depends on event variation
return "low" # likely overfitted to a specific command linedef confidence(fired: int, runs: int) -> str:
rate = fired / runs
if runs < 5:
return "insufficient-runs"
if rate == 1.0:
return "high"
if rate >= 0.8:
return "medium" # review: rule depends on event variation
return "low" # likely overfitted to a specific command lineThe thresholds above are illustrative; calibrate them on your own rules. The useful part is what a medium or low score means. A rule that fires on 3 of 5 generated variants is probably matching a specific string rather than a behavior. That result is a tuning signal, and it feeds the detection-debt metric from Part 1.
6. Architecture: Where the Oracle Plugs In
PR opened (detections/*.yml changed)
β
βΌ
CI: lint β validate metadata β sigma convert -t splunk
β
βΌ
CI: extract attack.tNNNN tags ββββββββββββββ
β β
βΌ βΌ
Oracle service Regression selector
provider: replay | generative (rules sharing CIM data model)
β β
βΌ β
HEC β index=synthetic_attacks ββββββββββββββ
β
βΌ
Splunk REST: run compiled SPL as ad hoc search job
β
βΌ
Verdict + confidence β PR check status
β
βΌ
Required human reviewer β promotePR opened (detections/*.yml changed)
β
βΌ
CI: lint β validate metadata β sigma convert -t splunk
β
βΌ
CI: extract attack.tNNNN tags ββββββββββββββ
β β
βΌ βΌ
Oracle service Regression selector
provider: replay | generative (rules sharing CIM data model)
β β
βΌ β
HEC β index=synthetic_attacks ββββββββββββββ
β
βΌ
Splunk REST: run compiled SPL as ad hoc search job
β
βΌ
Verdict + confidence β PR check status
β
βΌ
Required human reviewer β promoteImplementation notes, all configuration of the oracle rather than manual steps:
- Isolation. A dedicated index, a HEC token scoped only to it, and production correlation searches excluded from it, exactly as in the original article. The CI service account should hold no permissions on production indexes.
- Ad hoc search instead of schedule. ES correlation searches run on a schedule. Waiting for that schedule inside a PR check is slow and flaky. The oracle submits the compiled SPL directly as a search job over
index=synthetic_attackswith a bounded time range, then polls for completion. This validates the detection logic, not the scheduling and alert-action configuration, which needs its own test. - Time alignment. Generated events need timestamps inside the search window. Stamp them at ingestion time and keep any intra-sequence deltas the rule depends on, such as a threshold within a time window.
- Cleanup. Tag each run with a run ID field and delete or expire test events afterward, so one PR's data cannot satisfy another PR's rule.
7. What This Model Does Not Solve
- Semantic correctness is not realism. An event can satisfy the CIM fields and the rule logic while looking nothing like what a real endpoint emits. The oracle proves the rule matches the synthetic event, not that it matches the attack in your environment. Periodic comparison against real telemetry or a lab run remains necessary.
- Self-confirming tests. When the same model family writes the rule and generates the test data, they can share the same blind spot. Part 3 returns to this.
- Scheduling and alert actions are untested. Ad hoc validation covers detection logic only.
- Cost and latency. N generations per technique per PR adds up. Replay as the default provider, with generative synthesis reserved for techniques without datasets, keeps the pipeline fast.
- Benign coverage is thin. A handful of negative cases does not estimate production false-positive rate. Replaying a sample of real benign traffic in staging remains the stronger check.
8. Toward Part 3
With a defined input and output contract, the oracle can be called by anything, including something that is not a human. An agent that drafts a Sigma rule can submit it, read the verdict, and revise before opening a PR. Part 3 builds that loop and examines where it breaks, starting with the self-confirmation risk above.
9. Useful and Necessary Resources
Core tooling
- Sigma / pySigma / sigma-cli:
github.com/SigmaHQ/sigma - Splunk Attack Data:
github.com/splunk/attack_data - OTRF Security Datasets:
github.com/OTRF/Security-Datasets - Splunk Attack Range:
github.com/splunk/attack_range - Atomic Red Team:
github.com/redcanaryco/atomic-red-team - Splunk Security Content (ESCU) / contentctl:
github.com/splunk/security_content
Reference frameworks
- MITRE ATT&CK:
attack.mitre.org - Splunk CIM:
docs.splunk.com/Documentation/CIM
Reference reading
- Microsoft Security Blog, Accelerating detection engineering using AI-assisted synthetic attack logs generation, May 2026
- Huang et al., SAGA: Synthetic Audit Log Generation for APT Campaigns, arXiv:2411.13138
This article describes a detection engineering methodology for internal use. Test and production pipelines should be fully isolated, with promotion gated by a required human reviewer.
A.I. support for text and images
Detection as Code for Splunk: A CI/CD Approach to Correlation Rule Engineering Part 1 of a 3-part series on modern detection engineering. Part 2 connects this pipeline to Synthetic Attack Logβ¦
Synthetic Attack Log Generation for Splunk: A Detection Engineering Approach 1. Introduction and Scope