Measured, not asserted
Data quality
Every number below is read from a committed engine artifact — this page cannot say anything the runs didn't measure. Each stat links its reproducible evidence; each figure carries the date it was measured. The standing promises behind these guards live in DATA_SLA.md.
North star — full-surface audit ok-rate
Hundreds of cold, natural probes across every retrieval surface, graded against ground truth. The one number the whole engine system optimizes.
SCF funding cross-check
No project overstates or understates SCF membership vs the fund's own directory.
- 317 matched records checked against 547 SCF projects
- 271 per-round claims verified
- 13 round-level overclaims found by this run (queued as its fix wave)
Contract honesty probe
Documented params do something, undocumented values are rejected — the contract a stranger hits behaves as written.
- 1 silent param(s), 4 invalid-accepted param(s) at baseline (spec 1.7.16)
- Re-run on every deploy (post-deploy-eval workflow)
Code-depth calibration
Repo depth grades agree with independent code analysis where both exist.
- 0 disagreements
- 80 sampled repos had no independent index to compare (small-n baseline — grows as coverage does)
Consumer interlock (Raven)
The #1 consumer's discovery index tracks our contract — checked from OUR side too, with grace for their re-baseline cadence.
- 2 op(s) lagging within the 10-day re-baseline grace window (expected)
- 0 op(s) missing beyond grace
- contract 1.8.14 at measurement
Recall floors (Engine A)
Generated known-item probes per bucket stay above their red-line floors — recall can't silently erode.
- R-SYM at 64.3% vs floor 75% — open red
Real-demand OK-rate (Engine D)
The queries real consumers actually sent are replayed live — a miss on real demand outranks any synthetic finding.
- 5,313 real-consumer calls in the 14-day window
- 30 queries missing today — the standing fix queue
Golden retrieval eval
Known-true questions (answer key derived from the canonical directory) keep passing after every ship.
- full pass
- 2 N/A by design (live-source questions the static corpus doesn't answer)
- re-run on every production deploy + weekly
Corpus hygiene (S5–S8)
The research corpus stays clean — junk URLs, broken titles, staleness and mirror drift are swept weekly.
- s5 junk urls: 1 flagged — queued
- s6 bad titles: 69 flagged — queued
- s8 mirrors: 127 flagged — queued
Consumer path (through Raven)
The canonical questions answer correctly through the REAL Raven gateway — routing, our op, the response envelope and coaching — not just our direct API.
- every gradeable golden question answers correctly through Raven (44 checked)
Improvement ledger
Every quality detector's findings land in one tracked backlog by surface. A backlog is fine; a HIGH-severity finding neglected past 30 days is the failure — that's this row's red line.
- open by surface: retrieval 246 · directory 23 · contract 14 · consumer 14 · corpus 6 · code 4 · scf 0
- 404 tracked · 0 in-wave · 7 verified · 90 auto-cleared (detector stopped flagging)
- no high-severity finding neglected past 30 days
- ⚠ raven-routing (25d, 14 open), raven-loop (25d, 1 open) — evidence past the 10d window; 15 open finding(s) unconfirmed