Every malicious-package record in the bundled intel, across npm, PyPI, NuGet, RubyGems, Cargo, Go, Maven, Composer and VS Code, planted in a lockfile and scanned: 287,899 checks, none missed.
bench/malicious_records.pyMeasured against code nobody here wrote.
Every number on this page is produced by a script in the open-source repository, re-run for release 0.6.0, with the method beside it. The big ones are whole corpora, every item scanned: not a sample, and not a demo.
Every item, not a sample.
A number is only as good as what it was measured on. These runs scan every record and every file of their corpus, and each defect they found was fixed with a test before the number was written down.
The lockfiles of the most-downloaded projects on 14 registries, compared package by package. 98.7% of lockfiles agree on average, 12 of 14 registries at 98% or above, and every gap was read: each Cordon defect it found is fixed.
bench/parse_agreement.py3,941 of 4,034 attacks, each on its rule's scan path; 80.6% of the published evasions, read offline for intent; 92.4% of the catalogue's benign near-misses left clean.
bench/atr_bench.pyThe same samples, two tools.
GuardDog is the open-source malware scanner most teams reach for. Both ran over the same draw of real malware and the same popular packages, inside Docker with the network off.
Five questions, each against real code.
Malware from DataDog's public dataset and malregistry, every known-malicious record, the most-downloaded packages, real lockfiles and agent configs, and the infrastructure the cloud vendors publish as correct.
- Known-malicious records, by name and version100%249,646 records
- Real malicious packages detected (79.7% by the code alone)94.2%39,328 (DataDog, malregistry)
- The same malware, Cordon against GuardDog95.2% vs 85.5%498, a fixed-seed draw
- Popular packages wrongly blocked1.6% vs GuardDog's 16.8%top 1,000 PyPI + 1,000 npm
- Widely used open-source repositories passing the gate (0.4.0 run)85.4%1,427
- Packages in real lockfiles, against Trivy99.6%1,687 lockfiles, 14 registries
- Operating-system packages in 175 container images, against Syft99.9%21,758 packages (deb, apk, rpm, pacman, portage)
- Every package in those images, against Syft95.1% mean175 images, after stated exclusions
- CVEs, against Trivy and OSV-Scanner98.4%, every gap explained100 lockfiles
- Attacks, each on its rule's scan path97.7%4,034 cases
- Published evasions, read offline for intent80.6%289 cases
- Benign near-misses left clean92.4%4,369 cases
- Agent configs in real repositories blocked1.1%, each correct372 repositories
- Reference infrastructure, as its vendors publish it (0.4.0 run)2,560 findings, 795 blocking13 repositories, 20,310 files
Detection is not one number.
Whether a file attacks you, whether your graph resolves a named release, and whether an attack shape is covered at all need different evidence.
- Malicious npm packages (DataDog)93.7%25,766 samples · 75.8% by the code alone
- Malicious PyPI packages (DataDog)93.4%2,502 samples · 86.3% by the code alone
- Malicious AI-agent skills and IDE extensions (DataDog)35.6%326 samples · 35.6% by the code alone
- Malicious packages (malregistry)97.4%10,734 samples · 88.9% by the code alone
- All94.2%39,328 samples · 79.7% by the code alone
Public sample sets exist for npm and PyPI, so content is measured there; every ecosystem is measured by record, on the right.
- npm100%260,391 checked
- PyPI100%17,130 checked
- NuGet100%5,215 checked
- RubyGems100%5,047 checked
- VS Code extensions100%69 checked
- Cargo100%22 checked
- Go100%20 checked
- Maven100%3 checked
- Composer100%1 checked
- Git100%1 checked
Each record is planted as a pinned dependency in its ecosystem's own lockfile and scanned. Records whose only listed version is npm's empty takedown placeholder are counted and set aside: there is nothing malicious left to install.
Every attack shape, as a regression test.
The typosquat miss is deliberate: one character appended to a short name is how ecosystems name companion packages (vuex, reacts), so that shape is excused by name and reported by a different route.
- Install-hook attacks5 samples100%
- Credential exfiltration9 samples100%
- Obfuscated payloads7 samples100%
- Download-and-execute payloads11 samples100%
- Persistence mechanisms1 sample100%
- Supply-chain integrity4 samples100%
- CI/CD pipeline attacks2 samples100%
- Infrastructure misconfiguration4 samples100%
- Dependency confusion3 samples100%
- Typosquatting20 samples95%
Where the 5.9% goes.
Content detection reads what a package does. A package that does nothing yet, or ships only a compiled file, has nothing to read.
About 2 in 5 content misses are almost-empty packages: a bare package.json, a placeholder, a researcher's proof of concept. There is no behaviour in them to read. They are caught by name: all 249,646 known records are.
About 1 in 5 are a prebuilt binary and nothing else. Also caught by name, and a binary under a source name is its own finding.
The rest is the long tail, fixed shape by shape, each fix measured against the malware and the false-alarm corpora first.
Runtime-only behaviour is out of reach by construction: a payload decoded from a network response does not exist until the code runs. Cordon never runs it; the separate, opt-in sandbox is what observes that, now with a second install under a clock moved 400 days ahead.
Agent attacks, on the path each rule reads.
The Agent Threat Rules catalogue ships attack, evasion and benign cases for every rule. Cordon runs each on the scan path the catalogue gives it, then on 372 real repositories it was never tuned on.
docker build -f bench/Dockerfile -t cordon-bench:dev .docker volume create cordon-bench-datadocker run --rm -v cordon-bench-data:/data --entrypoint python cordon-bench:dev /bench/fetch.py all --count 15000docker run --rm --network none -v cordon-bench-data:/data:ro cordon-bench:dev malware benigndocker run --rm --entrypoint python cordon-bench:dev /bench/malicious_records.pydocker run --rm -v cordon-bench-data:/data --entrypoint python cordon-bench:dev /bench/fetch_lockfiles.pydocker run --rm --network none -v cordon-bench-data:/data:ro --entrypoint python cordon-bench:dev /bench/parse_agreement.py
The full method, per ecosystem and per technique: docs/08-ACCURACY.md at v0.6.0.