Measured, not asserted
Data quality
Every number here is read from a committed engine artifact, so this page cannot say anything the runs did not measure. Each figure links its evidence and carries the date it was taken. The standing promises live in DATA_SLA.md.
Verdict: where this data stands right now
Derived from the guard states below, not written by hand.
At target = passing on fresh evidence. Below target = measured, with an open work queue. Needs re-measure = evidence older than its own window. Open findings are ours and still reproducing; waiting on upstream is carried, not ours to fix.
11 of those have been waiting longer than 20 days — the oldest for 71. A re-baseline that has not happened in 71 days is not a wait any more.
Safe to rely on
value · days since measured
Row coverage: this page reads all 1133 project rows from the unranked listing, a census, so no row hides by being hard to retrieve.
Below target, being worked
A stale reading is a reading of the past, not the present — it is never counted as passing.
Progress against the quality plan
Phase status is read from QUALITY.md itself, a phase cannot show green here without being green there. Remaining work is shown in the same weight as completed work.
evidence ›evidence ⌄
evidence ›evidence ⌄
evidence ›evidence ⌄
evidence ›evidence ⌄
evidence ›evidence ⌄
evidence ›evidence ⌄
Reading this as an agent?
Every number on this page is served as JSON: the verdict block first, then the north star with its age, per-operation contract state, known limitations, the gap matrix with real identifiers, the miss funnel, consumer findings, guard state and the trend history. No parameters, no key. Cached one hour and served stale up to a day while revalidating, so read meta.measuredAt, not the clock.
North star: full-surface audit ok-rate
Hundreds of cold, natural probes across every retrieval surface, graded against ground truth. The one number the whole engine system optimizes.
Reading note: the 9 points were measured over different probe counts (597, 648, 198, 2164, 2235, 2229, 2247), so they do not form a comparable trend line
Known limitations: read before relying on this data
Derived from the measurements below, not written by hand: if a number improves, the entry changes or disappears. Each says what to do instead.
46% of lifecycle statuses rest on the weakest honest bases: a page answered (site-liveness) or a value inherited from a source (source-inherited).
Instead: Weigh statusBasis and statusAsOf on every row; treat human-verified and onchain-activity as the strong tiers, and verify a Live claim against the row's statusSourceUrl before repeating it.
Some rows carry no type, so an exact ?type= enumeration cannot see them even when the project belongs to that vertical.
Instead: For discovery, combine ?type= with a q= search; an empty typed result is a statement about our tagging, not about the ecosystem.
Curated, dated repo facts (knowledgeNotes) exist on a minority of indexed repositories.
Instead: Absence of notes is absence of curation, never evidence about the repo; fall back to codeVerified and activity fields.
Open findings by surface. A work queue, not an outage
An agent can fetch all of this, limitations, surface health, guard state, the score definitions and the trend, from GET /api/quality, so trust calibration does not require reading a webpage.
Gap matrix: what is missing, by entity and field
One row per hole. Counts are samples with their denominator, and every row carries real identifiers so the gap can be worked or independently checked.
Without a verified contract join we can say code exists, never that it is USED on mainnet. Denominator is the expected tier only — 5380 further low-signal deployable repos (demos/tutorials/experiments) are deliberately excluded: absence of a mainnet join there is expected-normal, and sweeping them in overstated this gap ~10×. Closed by: Contract attribution during the scan wave. Absence is absence of a join, never proof of disuse.
e.g. fazzatti/colibri, stellar/stellar-horizon, Shamba-Records-Limited/microvault, hyperledger-solang/solang, Phoenix-Protocol-Group/phoenix-contracts, stellar/sep45-reference
Curated dated facts are what let an agent cite a repo claim. Without them only raw scan fields are available. Closed by: The repo-intel enrich pass, which writes dated notes with sources.
e.g. luanlabs/fluxity-interface, Stellar-Light/stellar-pay, SoroWill/sorowill-sdk, Certora/CertoraProver, Vellar-Wallet/vellar-sdk, ravendevteam/alternet
sls-079: 'Live' says a product operates for users somewhere; it does NOT say which network it is deployed on. Of the 915 unknown rows, 201 are on-chain product types where the question applies and an agent will ask it; the other 714 are SDKs, wallets, security and analytics rows where unknown is the honest answer, not a gap. Closed by: Evidence only: a verified mainnet contract join, an on-chain activity reading, or a human-verified operator artifact (DEPLOYMENT_VERIFIED).
e.g. wisdomtree, redswan, spacewalk
site-liveness means only that a page answered. source-inherited means the value came from elsewhere. Neither is observation of the product. Of the weak rows, 6 hold an on-chain footprint (an issued asset, a joined contract, or a known deployment) and can earn onchain-activity from evidence; the other 474 are app-only and can only move through human verification — the ceiling of this row is people, not lanes. Closed by: Human verification with a receipt, an on-chain activity reading, or an operator announcement.
e.g. yellow-card, scorechain, solar-wallet, infstones
No depth reading means the repo was never scanned for real implementation signal. Closed by: A scan wave pass over the unscanned tail.
e.g. Dione-b/stellar-start, thiagocbalducci/neon-wallet-desktop, adletgamer/pakta, FernandoMay/stellar-energy-assets, karagozemin/stellar-drips, ManuelJG1999/Thalos
An untyped row is invisible to exact ?type= enumeration and to the gaps axis, even when it belongs to that vertical. Closed by: Add the type via the curation TYPE_ADD pass, with the row's own description as evidence.
e.g. tilt-pay, stellar-vrf, rail402, pigfi, lendwise, lendoor
A lifecycle claim with no source cannot be re-checked by a caller. It is our assertion, not evidence. Closed by: Curate a dated source URL, or downgrade the basis to match the evidence we actually have.
Agent lanes: autonomy earned, not assumed
A lane is a bounded, evidence-only job an agent runs on schedule. Advancement is measured — a lane earns auto-merge only after consecutive weeks where a human reviewed and changed nothing.
Weeks are counted only from successful scheduled runs, and the counter is derived daily from the lane's live write-set diffed against the committed snapshot — a quiet failure reads as a red week, never a clean one. A stamp a human upgrades to human-verified stays clean; a stamp a human removes or changes resets the count to zero.
Lane autonomy — intervention-free weeks
Every workflow in this repo that can write to production data, and the weeks each has earned toward running its execute unattended. Built by reading the workflow files, counted from GitHub's own run history, and reset by the intervention log — no lane's number is asserted here.
A week counts only when the lane executed and nothing it wrote was corrected.
| Lane | Cadence | Runs 8w | Weeks | Last correction |
|---|---|---|---|---|
| check-links | cron 23 2 * * * | 31 / 11 | 8 | — |
| enrich-onchain | cron 0 6 * * 2 | 8 / 15 | 8 | — |
| enrich-tvl | cron 0 5 * * 1 | 9 / 1 | 8 | — |
| refresh-research-corpus | cron 0 6 * * * | 48 / 22 | 8 | — |
| scan-repo-code | cron 17 */2 * * * · cron 43 3 1 * * | 298 / 157 | 8 | — |
| sync-lumenloop | cron 0 7 * * * | 51 / 2 | 8 | — |
| aggregate-feedback | cron 40 3 * * * | 45 / 2 | 7 | — |
| backfill-triage-tags | cron 0 4 * * * | 43 / 2 | 7 | — |
| enrich-builders | cron 30 6 * * * | 42 / 3 | 6 | — |
| enrich-repo-activity | cron 20 4 * * * | 41 / 4 | 6 | — |
| enrich-repos | cron 0 8 * * 1 | 7 / 78 | 6 | — |
| refresh-stablecoins | cron 25 */6 * * * | 134 / 49 | 6 | — |
| enrich-builder-repos | cron 0 5 * * 2 | 5 / 7 | 5 | — |
| raven-eval-parity | cron 30 7 * * * | 24 / 19 | 5 | — |
| upgrade-status-basis | cron 23 4 * * * | 29 / 14 | 5 | — |
| backfill-knowledge-notes | cron 23 3 * * * | 26 / 34 | 3 | 2026-09-05 · The nightly step was gated to workflow_dispatch, so the 3 scheduled runs since the cron was added on 2026-09-02 executed nothing (18 completed runs in total, most manual dispatches; corrected 2026-09-05 from the runs' job steps); the gate was removed so schedule runs write. Found by the lane-autonomy counter's first run. |
| generated-recall | cron 0 7 * * 0 | 4 / 1 | 3 | — |
| refresh-rwa | cron 37 */6 * * * | 90 / 2 | 3 | 2026-09-05 · The sls-023 close-out on issue #494 was overclaimed — 11 of the finding's 61 rows were treated as closing the whole probe. Recorded as a violated lesson ("verify before advertise (#494, twice)"); the close-out itself has not yet been reframed upstream, so this stays open as a correction against the lane that now owns those rows. |
| verify-published-packages | cron 40 5 * * 2 | 3 / 5 | 3 | — |
| award-reconcile | cron 30 5 * * * | 11 / 14 | 2 | — |
| basis-from-deployment | cron 30 8 * * 2 | 2 / 7 | 2 | — |
| curate-projects | cron 30 7 * * 2 | 3 / 171 | 2 | 2026-09-05 · A July STATUS_FIX entry for blend (Live→Live, basis site-liveness) re-stamped site-liveness on every execute, erasing the onchain-activity the basis lane awards; 16 entries carry a weak basis. Fixed: a weak curated basis never overwrites a strong lane-earned one when the status is unchanged. |
| basis-from-onchain | cron 0 9 * * 2 | 1 / 13 | 1 | — |
| basis-from-product | cron 0 10 * * 2 | 1 / 8 | 1 | 2026-09-04 · The lane's writes were correct, its read-back was not: STRONG_BASES still listed only three tiers, so the 58 rows this lane awarded `product-integration` were counted as weak and the board read the completed work as failed. |
| basis-from-repo | cron 30 9 * * 2 | 1 / 4 | 1 | 2026-09-04 · Same defect, same PR: the 115 rows this lane awarded `repo-activity` were excluded from the strong-basis metric, so weakLiveRows read 790 when it was 617. |
| upgrade-basis-onchain | cron 0 9 * * 2 | 1 / 12 | 1 | — |
| apply-content-pass | dispatch | 0 / 0 | 0 | — |
| apply-status-batch | dispatch | 0 / 0 | 0 | — |
| apply-wallet-availability | dispatch | 0 / 6 | 0 | — |
| award-import | dispatch | 0 / 27 | 0 | — |
| award-publish | dispatch | 0 / 7 | 0 | — |
| award-round | dispatch | 0 / 54 | 0 | — |
| award-test-setup | dispatch | 0 / 31 | 0 | — |
| backfill-code-domains | dispatch | 0 / 5 | 0 | — |
| backfill-code-tier | dispatch | 0 / 0 | 0 | — |
| backfill-contract-basis | dispatch | 0 / 2 | 0 | — |
| backfill-lifecycle | dispatch | 0 / 0 | 0 | — |
| backfill-project-github | dispatch | 0 / 2 | 0 | — |
| backfill-provenance | dispatch | 0 / 2 | 0 | — |
| backfill-version-status | dispatch | 0 / 2 | 0 | — |
| curate-partners | dispatch | 0 / 6 | 0 | — |
| db-space | dispatch | 0 / 0 | 0 | — |
| dedup-projects | dispatch | 0 / 3 | 0 | 2026-09-05 · Its 11 Draft hides were overwritten 30 minutes later by curate-projects' DUPE_MERGES step, which forces status=Inactive on every shadow (scripts/data/curate-projects.ts, 'if (dupe.status !== "Inactive")'). Two lanes own one field with different verdicts; the July fold (Inactive + canonicalSlug) won silently. Rows stay hidden from search and the board either way; the wrong status on duplicates is the open class in improvements/drafts/2026-09-05-inactive-site-liveness-triage.md. |
| dedupe-repos | dispatch | 0 / 5 | 0 | — |
| embed-projects | dispatch | 0 / 0 | 0 | — |
| enrich-entities | dispatch | 0 / 9 | 0 | — |
| enrich-partner-onchain | dispatch | 0 / 0 | 0 | — |
| enrich-partners | dispatch | 0 / 3 | 0 | — |
| enrich-scf | dispatch | 0 / 34 | 0 | — |
| fix-dev-docs-titles | dispatch | 0 / 1 | 0 | — |
| fix-research-titles | dispatch | 0 / 5 | 0 | — |
| fix-scf-rounds | dispatch | 0 / 4 | 0 | — |
| fix-zenex | dispatch | 0 / 0 | 0 | — |
| gone-repos | cron 40 6 * * * | 0 / 1 | 0 | — |
| import-replit-stablecoin-history | dispatch | 0 / 2 | 0 | — |
| ingest-dora-evals | dispatch | 0 / 0 | 0 | — |
| ingest-ec-taxonomy | dispatch | 0 / 9 | 0 | — |
| link-canonical-slug | dispatch | 0 / 0 | 0 | — |
| link-partner-projects | dispatch | 0 / 0 | 0 | — |
| mark-defunct | dispatch | 0 / 0 | 0 | — |
| mark-inactive-projects | dispatch | 0 / 2 | 0 | — |
| migrate-research-url-host | dispatch | 0 / 0 | 0 | — |
| patch-rozo | dispatch | 0 / 0 | 0 | — |
| prune-denied-repos | dispatch | 0 / 2 | 0 | — |
| regrade-repos | dispatch | 0 / 25 | 0 | — |
| repair-lane | cron 47 8 * * * | 0 / 2 | 0 | — |
| seed-blog-posts | dispatch | 0 / 3 | 0 | — |
| seed-i3-mock | dispatch | 0 / 0 | 0 | — |
| seed-partners | dispatch | 0 / 0 | 0 | — |
| seed-summit-sp-winners | dispatch | 0 / 4 | 0 | — |
| set-project-status | dispatch | 0 / 0 | 0 | — |
| set-prominence | dispatch | 0 / 0 | 0 | — |
| set-types | dispatch | 0 / 0 | 0 | — |
| tansu-anchor | dispatch | 0 / 12 | 0 | — |
| task-lane | cron 17 9 * * 2 · cron 17 9 * * 3 · cron 17 9 * * 4 · cron 17 9 * * 5 | 1 / 0 | 0 | — |
| vector-index | dispatch | 0 / 2 | 0 | — |
how a week is counted ›how a week is counted ⌄
Elapsed time earns nothing. The weeks must be consecutive and must run up to this one, so a lane nobody has run sits at zero however long it has been quiet, and four scattered good weeks are not four clean weeks. Only runs the lane started ITSELF count — a hand-dispatched execute is a person operating the lane, and is reported here rather than counted. The run counts are runs, not writes: which of them wrote is what the week count is proven from. Every counted execute is proven from that run's own job steps, never from today's copy of the workflow file: a step that was skipped moved nothing, whatever the file says now. A “+” marks a floor: GitHub does not expose a run's commands, so a step the author named and that actually ran could not be classified. Corrections live in improvements/lanes/interventions.json, appended by the same PR that makes the correction.
Findings: what the engines caught
Every detector writes here. This is the work queue, not a score.
The three states are disjoint and sum to 741. Cleared is NOT confirmation the fix works; only verified means it was deliberately re-probed after a fix.
Both are still-open rows across all three buckets above (open, waiting on upstream, and the refresh queue), split by what each finding IS rather than whose turn it is. A world finding is repaired by curation and is evidence the detectors do their job; an instrument finding is ours. Lifetime: 268 world, 473 instrument. The mapping is one table per detector and failure mode in src/lib/improvement-ledger.ts; an unmapped pair counts as instrument, never as the product working.
A recurrence is a NEW finding on a surface-and-failure-mode pair we had already closed on silence — the detector went quiet and nobody re-probed. It is the closure rule's real question: did we close without repairing, and did the same kind of failure come back? Steady state is that rate at zero. Reopened counts the exact-id version and is a lower bound, because re-clearing a reopened finding erases its stamp. These numbers are expected to start ugly; publishing them is the point.
For context, the repeat-class rate — a finding whose §0 class (identity, taxonomy coverage, contract completeness…) already had any prior finding — is 99.5% over 30 days across 7 classes. With classes that broad it cannot fall, so it is reported as context rather than steered by.
By failure mode and state
How long the open ones have been open
A tall bar on the right is the treadmill this page exists to end: detection outrunning remediation. An empty right bar just after a stale-findings sweep means old entries were re-probed and cleared, not remediated.
Recently cleared
Consumer findings from Raven
Defects filed against this service by stellar-raven, its largest agent consumer, from that project's own evaluation battery. These carry more signal than our internal detectors because the answer key is not ours.
Their answer key, by status
A record can read reported-upstream on their side while our linked issue is closed: we shipped the fix and their re-verification has not run yet. Neither status is allowed to speak for the other.
Defect flow: detector to surface to outcome
Every finding in the ledger traced through the system: which detector caught it, which surface it lives on, and whether it closed. Ribbon thickness is the count; whole-ledger, not a sample.
Read left to right: a detector produces findings, they land on a surface, and they end Cleared, Open, or Verified. A fat ribbon into Open is a surface carrying real debt; a fat ribbon into Cleared is a detector whose class has been closed.
Where open recall misses die
Each open recall finding replayed live and classified at the FIRST stage that fails - mutually exclusive classes with different owners, not a funnel or a sequence.
0 of 0 open recall findings replayed
Row quality: the evidence behind each record
Every project row scores on five facts we either hold or don't: a provenance basis, a date, a source URL, a type, and a link. A low score names exactly what is missing.
Status provenance, strongest evidence first
Deployment fact (sls-079): 915 unknown · 87 mainnet · 2 testnetiWhich network a product is deployed on, as a separate fact from lifecycle status. Populated ONLY from evidence (verified mainnet contract joins, on-chain readings, human-verified operator artifacts); unknown is the honest default and a work queue, never a score. The gap matrix carries the prominent rows to work first.
Most rows rest on site-liveness - a page answered. That is the weakest honest basis we serve, and moving rows up this ramp is the standing data job.
What is missing, across all rows
Every row as one dot. The marked region is the curation queue.
Curation queue: rows whose evidence is thinnest, prominent firstiSorted by evidence score ascending, then by curated prominence, so the most-seen thin rows surface first. Each line names exactly which of the five facts is missing.
The code index: what it holds, how deeply we know it
2,920 curated repos (claimed by a project or a tracked builder) plus a 10,018-row Electric Capital tail indexed for completeness. The charts read over the curated index only - mixing the tail in made the curated index look unscanned when it is not.
Repo score distribution, curated indexirepoScore (0-100) grades freshness, traction and builder authority. A long low tail is EXPECTED in an open ecosystem - hackathon one-offs and early experiments are real code references worth indexing; the score is what keeps them ranked below production repos.
Commit activity, curated indexiDerived from each repo's last commit: active (<=90d), slowing (<=1y), dormant (older), archived, or unknown (no commit date held - not knowing is its own state, never counted as dormant).
Languages
How deeply we know each layeriThree different jobs with three different denominators. Depth scanning aims at the whole curated index. Knowledge notes are a hand-curated research layer being built over the highest-scored repos - a small number is early progress, not missing homework. Mainnet joins are deliberately strict: only a verified on-chain attribution counts, so the number grows slowly and every unit of it is proof.
Curated repos with a code-depth reading (entry files, symbols, SDK usage). The scan waves aim at all of them.
Hand-curated dated facts with sources, written over the top-scored pool one repo at a time. An enrichment layer under construction, newest additions first.
Deployable contracts with a PROVEN on-chain attribution. Strict by design: absence is absence of a join, never proof of disuse.
Notes pool, honestly split: of 373 pool repos, 342 carry dated facts, 29 were examined and yielded nothing durable (judged, recorded internally), and 2 are still unexamined. A judged repo is not a gap.
Plus the Electric Capital tail: 7,948 of 10,018 rows scanned opportunistically as budget allows - indexed for completeness, no coverage target attached.
Highest-graded repos
Lessons and research
Every recurring defect class was written up when it was found, and every human-verified correction carries a committed receipt. These are the documents behind the numbers above.
Lesson write-upsiEach file records defects found on one day: what broke, the root cause, and the invariant or probe added so the class cannot silently return.
A lesson counts only when it became a check — a test, a guard script or a workflow that goes red when the class returns. Each file names its guard with a Guard: line; check-lessons-guarded verifies the file exists and fails on any that is unguarded, on every PR.
Correction receiptsiA human-verified status change commits its evidence: the URL fetched, the time, response identity headers, and the exact markers on the page that decided the verdict. Re-run the capture to diff what a page says now against what it said then.
Audits
Trends
Daily history appended by the eval pipeline and committed, red days included. Battery probe counts rotate with the daily banks, so the pass line moves by design; the failure line and the ratchets are the signal.
SCF funding cross-check
No project overstates or understates SCF membership, at the project level OR the round level, against the fund's own directory.
- 340 matched records checked against 531 SCF projects
- 0 overstated / 0 understated at project level
- 0 round-level overclaims across 328 verified claims
Contract honesty probe (Engine E)
Documented params do something, undocumented values are rejected. The contract a stranger hits behaves as written.
- 0 param(s) documented but silently ignored
- 0 param(s) accepting values the spec forbids
- spec 1.9.54 at measurement; probed on every deploy, evidence committed weekly by engine-c-health
Curated canonical repos resolve
Every repo we call authoritative is indexed and carries code signals — a curated name that matches no row silently degrades the query it was written for.
- 0 absent (curated name matches no row)
- 0 indexed but no code signals — invisible to code-evidence ranking and to the tier gate
Human-verified stamps still hold
Every human-verified packet stamp is re-probed weekly against its own deciding URL; a contradiction is a finding, not a silent stale claim.
- 2 could-not-check — a 403, a timeout or a client-rendered shell is a page we did not read, never a contradiction
- read-only: this guard files a finding for a human and never writes a status, because reading a marketing banner as a product state is the failure it exists to catch
Computed values reach a serving path
Every field our machinery computes is read by something that shapes an agent's answer — a value nothing consumes cannot change what anyone is told.
- a script or a test does not count as consumption — that is how codeProofTier passed for months while only a report called it
Guard lanes that actually run
Every automated lane in the repo has completed a real run recently — a guard that never executes is a promise, not a check.
- STANDING-SIGNAL: gone-repos.yml — 19 run(s) red at its declared signal step since its last green 17d ago — the lane works; what it measures has failed every run and nobody has acted (step "Probe")
- STANDING-SIGNAL: raven-eval-parity.yml — 3 run(s) red at its declared signal step since its last green 1d ago — the lane works; what it measures has failed every run and nobody has acted (step "Truth battery (guard D — rotating probes, curated answer key)")
- STANDING-SIGNAL: task-lane.yml — 7 run(s) red at its declared signal step since its last green 13d ago — the lane works; what it measures has failed every run and nobody has acted (step "Verify the end state")
- judged against the workflow file as it stands: runs from a since-edited or never-merged version are ignored, so a fixed lane stops being red
- a lane that exits 1 to report a finding is the guard working, and is not counted here
Type errors in scripts/
The scripts that write to the production database are type-checked, and the backlog of known errors only shrinks.
- identity is file + error code + message, without line numbers, so an unrelated edit above an error does not read as a regression
- a count of 0 here means the ratchet is finished and the guard can become a plain tsc gate
Served values a consumer can date
Every value we serve carries a date that covers IT — or says plainly that it cannot be dated. A value wearing a neighbour's timestamp is a wrong answer with a citation.
- explainRepo.answerAsOf is the pattern: NULL for a DeepWiki answer, because DeepWiki exposes no index date and inventing one would make an unknown look measured
- scoping cuts both ways: verifyClaim's confidence.ageDays dates the numbers inside confidence but NOT the root-level verdict beside it — so verifyClaim.verdict is carried as named debt, not excused; an admission only counts when the description speaks to datability itself ('no index date'), never bare 'null' or 'unknown'
- limit: a scope holding one date is treated as dating every value in it, so this catches the sharper shape only — a value with NO date in scope while other objects in the response carry dates
SCF-funded projects served
Every SCF-funded project is in the directory an agent searches, with funding read from SCF's own submission records.
- 10 SCF-funded projects the directory does not serve
- DefiLlama: 1 missing of 42 Stellar-listed
- absent entries are human-reviewed SEEDS, never bulk-created
- absent: open-rwa-vault-infrastructure-nis (round 45)
- absent: canary-protocol-y4l (round )
- absent: pods-dxl (round )
Code-depth calibration
Repo depth grades agree with independent code analysis where both exist — on a sample large enough for the rate to mean something.
- 0 disagreements
- two graders, recorded per row and never pooled silently: DeepWiki for the few answer-key repos it has indexed, Grok (agentic, reads the repo) for the rest
- the binding constraints are EXTERNAL: DeepWiki's index covers ~4 of the 86 long-tail answer-key repos, and the Grok balance exhausted mid-run (HTTP 402) after 9 verdicts — topping it up or DeepWiki indexing more of the key is what grows n
- sample meets the 20-repo floor
Consumer interlock (Raven)
The #1 consumer's discovery index tracks our contract, checked from OUR side too, with grace for their re-baseline cadence.
- 0 op(s) lagging within the 10-day re-baseline grace window (expected)
- 0 op(s) missing beyond grace
- 3 op(s) in the catalog that our own routing words do not reach — our fix
- 3 op(s) callable but absent from the consumer's discovery index past the 10-day grace window (getQualityReport, getRwaAssets, verifyClaim) — an agent that does not already know the name cannot find them, and no wording of ours can change that
- contract 1.9.54 at measurement
Recall floors (Engine A)
Generated known-item probes per bucket stay above their red-line floors, recall can't silently erode.
- all buckets above floor
Real-demand OK-rate (Engine D)
The queries real consumers actually sent are replayed live and keep answering at or above the committed floor. A miss on real demand outranks any synthetic finding.
- 64,185 real-consumer calls in the 14-day window
- 0 queries missing today, the standing fix queue
- floor 80% is a ratchet from a prior reading, it may only move up
Golden retrieval eval
Known-true questions (answer key derived from the canonical directory) keep passing after every ship.
- full pass
- 2 N/A by design (live-source questions the static corpus doesn't answer)
- re-run on every production deploy + weekly
Corpus hygiene (S5-S8)
The research corpus stays clean: junk URLs, broken titles, staleness and mirror drift are swept weekly, and a sweep that returns no reading counts as unknown, never as clean.
- s7 coverage: 2 source(s) structurally undateable (declared)
- all sweeps clean
Consumer path (through Raven)
The canonical questions answer correctly through the REAL Raven gateway: routing, our op, the response envelope and coaching, not just our direct API.
- every gradeable golden question answers correctly through Raven (47 checked)
Improvement ledger
Every quality detector's findings land in one tracked backlog by surface. A backlog is fine; a HIGH-severity finding neglected past 30 days is the failure, that's this row's red line.
- open by surface: code 8 · retrieval 0 · directory 0 · scf 0 · contract 0 · consumer 0 · corpus 0
- 741 tracked · 0 in-wave · 7 verified · 653 auto-cleared (detector stopped flagging)
- closure basis: 7 verified deliberately · 216 cleared on a live re-probe that PASSED · 437 cleared only because a detector stopped reporting
- that last group is the re-probe backlog, not a result: a detector going quiet is indistinguishable from a gap nobody asked about again. Spot-check 2026-08-31 — kutana, etesia and octopos sit in it, are still absent from the directory, and each carries SCF round badges.
- no high-severity finding neglected past 30 days
- ⚠ raven-routing (17d, 17 open), evidence past the 10d window; 17 open finding(s) unconfirmed
Response-shape opacity (build-enforced)
No silent opacity: every response object declares its shape. Grandfathered open maps are counted here and the count may only fall.
- 0 response objects still return an undeclared shape
- the ratchet fails the build if that number rises; it may only fall
- a NEW operation shipping without a declared shape fails the build outright
- scripts/contract/check-schema-opacity.ts, wired into contract:check
Match labelling (build-enforced)
Every q-taking operation declares HOW it matched, so a caller can tell an exact hit from a semantic neighbour. Enforced in CI.
- 4 q-taking operation(s) tracked, 0 unlabelled
- a new q-taking operation shipping without a match label fails the build
- scripts/contract/check-honesty-layer.ts, wired into contract:check