Measured, not asserted

Data quality

Every number below is read from a committed engine artifact, this page cannot say anything the runs didn't measure. Each stat links its reproducible evidence; each figure carries the date it was measured. The standing promises behind these guards live in DATA_SLA.md.

Verdict: where this data stands right now

Derived from the guard states below, not written by hand. A guard is at target only when its evidence is both passing and fresh; below-target rows carry their own work queue, and aged evidence counts as unmeasured, never as passing.

At target
17
Passing on fresh evidence
Below target
1
Measured, with an open work queue
Needs re-measure
0
Evidence older than its own window
Open findings
5
Still reproducing on the latest run

Safe to rely on

  • SCF funding cross-check · 0 · measured 2026-08-28
  • Contract honesty probe (Engine E) · 0 · measured 2026-08-28
  • Curated canonical repos resolve · 44/44 · measured 2026-08-31
  • Computed values reach a serving path · 9/9 · measured 2026-08-31
  • Type errors in scripts/ · 63 · measured 2026-08-31
  • Served values a consumer can date · 46/80 · measured 2026-08-31
  • SCF-funded projects served · 500/500 · measured 2026-09-01
  • Code-depth calibration · 100% (n=28) · measured 2026-07-10
  • Consumer interlock (Raven) · 29/33 · measured 2026-08-28
  • Recall floors (Engine A) · 8/8 · measured 2026-09-01
  • Real-demand OK-rate (Engine D) · 99% · measured 2026-09-01
  • Golden retrieval eval · 51/51 · measured 2026-09-04
  • Corpus hygiene (S5-S8) · 5/5 · measured 2026-09-01
  • Consumer path (through Raven) · 47/47 · measured 2026-08-28
  • Improvement ledger · 5 open · measured 2026-09-04
  • Response-shape opacity (build-enforced) · 0 · measured 2026-09-04
  • Match labelling (build-enforced) · 0 · measured 2026-09-04
  • Row coverage · this page reads all 1100 project rows from the unranked listing, a census, so no row hides by being hard to retrieve

Below target, being worked

  • Guard lanes that actually run is below target: 84/86 lanes, measured 2026-09-04

Progress against the quality plan

Phase status is read from QUALITY.md itself, a phase cannot show green here without being green there. Remaining work is shown in the same weight as completed work.

3 of 6 phases complete
P0Name the classes, lock the first invariantDoneEvidence: check-schema-opacity.ts (the 47 baselined open maps were paid down to ZERO in #1092; the ratchet now holds the floor at 0), QUALITY.md itself.
P1Honesty layer, eval integrity, the dashboardDoneEvidence: specs/honesty-baseline.json (debt 0), scripts/eval/eval-baselines.json, improvements/quality/*.json.
P2Entity truth: issuers, receipts, enumerations, dedupeDoneEvidence: battery slice G (enumeration integrity), slice H (verify grades itself), improvements/receipts/, improvements/quality/entities.json (repos.duplicateRows = 0), scripts/data/check-repo-dupes.ts in enrich-repos.yml + dedupe-repos.yml.
P3Earned autonomyIn progressShipped so far: the daily pipeline now rebuilds its own quality artifacts and commits them, and the stale-finding sweep re-probes the ledger instead of letting counts drift. 2026-08-28: the FIRST bounded agent lane is live — the deployment-evidence gap (sls-079) is worked mechanically by the weekly curation pass via the operator-toml chain (the project's own stellar.toml -> declared code+issuer -> confirmed on Horizon mainnet; full chain or abstain, basis labeled "operator-toml" so a machine stamp never impersonates a human one). The lane reproduces the 2026-08-28 hand-worked queue's mechanical half; judgment cases (operator docs, bundles) stay human. 2026-08-29: the closure rule's METRIC exists — repeat-class rate is computed from the ledger and published on /quality (first measurement: 30-day rate 100%, 168/168 new findings in already-seen classes; lifetime 98.7% across 6 classes + meta-eval. The treadmill, now with a number).Remaining: one lane is not a system — the gap matrix's other rows (typed, sourced, knowledge notes) still close by hand, and Stage 2 requires N intervention-free weeks before auto-merge opens for bounded lanes; the count starts now, at zero.
P4Basis strength at scaleIn progressShipped so far: 2026-08-31→09-01 — the basis-upgrade lanes exist and have run against the population: evidence A (asset movement deltas between two dated stellar.expert readings), B (DeFiLlama TVL ≤14d) and C (Horizon, same day — issuer payments asset-matched over 20 records, XLM-pair trade fallback), every probe trinary (hit / checked-empty / could-not-check reaches the summary and the exit code). 35 rows now rest on onchain-activity; a two-auditor pass (agent + Grok) reverted 4 uncorrected-probe upgrades, one of which re-earned its upgrade the same run under the corrected rule. Asset keys joined from the stablecoin registry (12) and operators' own stellar.toml (7) so deltas compound weekly. Death receipts: 42 stamped, 8 retracted after audit (a live-200 page cannot stand as "observed dead"), 2 re-stamped once their domains went hard-dead. Untyped 59→2 (both honest residuals). Weak share 86%→84% (830 of 984 served rows).Remaining: the done bar is weak bases under 50%. 144 weak rows have a website that never answered a successful check — their reason is now printed as an owner triage table (relink / Inactive / leave), and that is human work, not a lane. operator-announcement is a basis value on 3 rows: a corpus-announcement lane (dated SDF/operator launch posts → basis + receipt) is the unbuilt lever with the most headroom. The XLM-denominated channel deposit has no USD ceiling until a price source that path may depend on exists.
P5The knowledge layer consumers keep asking forIn progressShipped so far: 2026-09-01 — the curated knowledgeNotes registry grew 16→~29 repos, every note dated and source-cited, including the supersession facts consumers actually ask for as prose (stellar/go → go-stellar-sdk, the Horizon monorepo split, the js-sdk deprecation chain, protocol ceilings). explainRepo now answers from a dated note ahead of an undated DeepWiki walkthrough (answerSource: "knowledge-note", answerAsOf RFC 3339), matched by exact identifier or by hand-authored trigger phrases; the matcher was hijack-hardened (citation URLs and bare domains can no longer route a note). sls-080 closed on the consumer's own probe and independently re-verified by Raven on 2026-09-01.Remaining: supersededBy / deprecatedAt still exist nowhere as FIELDS — the facts live in note prose, which a consumer must read rather than join on. Contracts as first-class joined entities only where the P3 lane reached (11 of the 308 expected-tier repos). Builder/org identity is still thinner than project identity. 12,851 indexed repos carry no note — the long tail is by design, the curated pool is the floor that rises.

Reading this as an agent?

Every number on this page is served as JSON: the verdict block first, then the north star with its age, per-operation contract state, known limitations, the gap matrix with real identifiers, the miss funnel, consumer findings, guard state and the trend history. No parameters, no key. Cached one hour and served stale up to a day while revalidating, so read meta.measuredAt, not the clock.

GET /api/quality

North star: full-surface audit ok-rate

Hundreds of cold, natural probes across every retrieval surface, graded against ground truth. The one number the whole engine system optimizes.

Reading note: the 5 points were measured over different probe counts (597, 648, 198, 2164), so they do not form a comparable trend line

latest (2026-08-28)
99%
2133/2164 probes ok · target ≥85%

Known limitations: read before relying on this data

Derived from the measurements below, not written by hand: if a number improves, the entry changes or disappears. Each says what to do instead.

project status606 of 984 rows (census)

62% of lifecycle statuses rest on the weakest honest bases: a page answered (site-liveness) or a value inherited from a source (source-inherited).

Instead: Weigh statusBasis and statusAsOf on every row; treat human-verified and onchain-activity as the strong tiers, and verify a Live claim against the row's statusSourceUrl before repeating it.

project types2 of 984 sampled rows untyped

Some rows carry no type, so an exact ?type= enumeration cannot see them even when the project belongs to that vertical.

Instead: For discovery, combine ?type= with a q= search; an empty typed result is a statement about our tagging, not about the ecosystem.

repo knowledge notes151 of 408 rows in the curation pool (curated index, repoScore >= 50)

Curated, dated repo facts (knowledgeNotes) exist on a minority of indexed repositories.

Instead: Absence of notes is absence of curation, never evidence about the repo; fall back to codeVerified and activity fields.

Open findings by surface. A work queue, not an outage

consumer
5
code
1
retrieval
0
scf
0
contract
0
directory
0
corpus
0

An agent can fetch all of this, limitations, surface health, guard state, the score definitions and the trend, from GET /api/quality, so trust calibration does not require reading a webpage.

Gap matrix: what is missing, by entity and field

One row per hole. Counts are samples with their denominator, and every row carries real identifiers so the gap can be worked or independently checked.

repo mainnet contract join381 / 395 missing (96%)

Without a verified contract join we can say code exists, never that it is USED on mainnet. Denominator is the expected tier only — 5265 further low-signal deployable repos (demos/tutorials/experiments) are deliberately excluded: absence of a mainnet join there is expected-normal, and sweeping them in overstated this gap ~10×. Closed by: Contract attribution during the scan wave. Absence is absence of a join, never proof of disuse.

e.g. Andy00L/x402-autopilot, Bosun-Josh121/clevercon, davidmaronio/StellarPay402, ritik4ever/lodestar, velikanghost/heekowave, 402md/agentcard

repo knowledgeNotes12633 / 13129 missing (96%)

Curated dated facts are what let an agent cite a repo claim. Without them only raw scan fields are available. Closed by: The repo-intel enrich pass, which writes dated notes with sources.

e.g. Andy00L/x402-autopilot, velikanghost/heekowave, thewoodfish/AgentCompute, 402md/agentcard, ashfrancis/chickenz, catmcgee/stellar-poker-cosnarks

project deployment evidence903 / 984 missing (92%)

sls-079: 'Live' says a product operates for users somewhere; it does NOT say which network it is deployed on. Of the 903 unknown rows, 199 are on-chain product types where the question applies and an agent will ask it; the other 704 are SDKs, wallets, security and analytics rows where unknown is the honest answer, not a gap. Closed by: Evidence only: a verified mainnet contract join, an on-chain activity reading, or a human-verified operator artifact (DEPLOYMENT_VERIFIED).

e.g. wisdomtree, redswan, spacewalk

project strongBasis617 / 984 missing (63%)

site-liveness means only that a page answered. source-inherited means the value came from elsewhere. Neither is observation of the product. Of the weak rows, 34 hold an on-chain footprint (an issued asset, a joined contract, or a known deployment) and can earn onchain-activity from evidence; the other 583 are app-only and can only move through human verification — the ceiling of this row is people, not lanes. Closed by: Human verification with a receipt, an on-chain activity reading, or an operator announcement.

e.g. kale, allbridge, ondo, quicknode, dia, axelar

repo code depth reading2282 / 13129 missing (17%)

No depth reading means the repo was never scanned for real implementation signal. Closed by: A scan wave pass over the unscanned tail.

e.g. leocagli/pontepay, leocagli/tralalero-contracts, KingFRANKHOOD/soroban-sql-sync, KB2410/ai-wallet-manager, KaruG1999/ocean_request, karagozemin/frontend-app

project sourced15 / 984 missing (2%)

A lifecycle claim with no source cannot be re-checked by a caller. It is our assertion, not evidence. Closed by: Curate a dated source URL, or downgrade the basis to match the evidence we actually have.

e.g. brale, ylds, mydatacoin, transfermole, the-blue-marble, scam-flagging-system

project typed2 / 984 missing (0%)

An untyped row is invisible to exact ?type= enumeration and to the gaps axis, even when it belongs to that vertical. Closed by: Add the type via the curation TYPE_ADD pass, with the row's own description as evidence.

e.g. boundless-bounties, deb

Agent lanes: autonomy earned, not assumed

A lane is a bounded, evidence-only job an agent runs on schedule. Advancement is measured — a lane earns auto-merge only after consecutive weeks where a human reviewed and changed nothing.

Lane: operator-toml
8 stamps
deployment facts, full-chain-or-abstain
Intervention-free weeks
0 / 2
to Stage 2 (auto-merge for bounded work)
Last run
cancelled
workflow_dispatch · 2026-09-02
Corrections
0
human had to fix a stamp — resets the counter

Weeks are counted only from successful scheduled runs, and the counter is derived daily from the lane's live write-set diffed against the committed snapshot — a quiet failure reads as a red week, never a clean one. A stamp a human upgrades to human-verified stays clean; a stamp a human removes or changes resets the count to zero.

Findings: what the engines caught

Every detector writes here. This is the work queue, not a score.

Open
5
Still reproducing
Cleared
514
stopped reproducing
Verified closed
7
Re-probed after the fix

The three states are disjoint and sum to 527. Cleared is NOT confirmation the fix works; only verified means it was deliberately re-probed after a fix.

Repeat-class rate (30d)
100%
235 of 235 new findings
Lifetime
98.9%
across 527 findings, 6 classes

A repeat is a finding whose class (identity, taxonomy coverage, contract completeness…) already had a prior finding — the measure of whether fixes land on the class or just the instance. Steady state is the 30-day rate at zero: new findings only ever open new classes. This number is expected to start ugly; publishing it is the point.

By failure mode and state

openclearedverifiedrecall-miss
coverage-gap
demand-miss
completeness-residual
api-drift
battery-coverage-weak
scf-round-overclaim
demand-routing-miss
routing-miss
fewermoreempty cell = zero

How long the open ones have been open

≤7d
2
8–30d
1
31–60d
3
>60d
0

A tall bar on the right is the treadmill this page exists to end: detection outrunning remediation. An empty right bar just after a stale-findings sweep means old entries were re-probed and cleared, not remediated.

Consumer findings from Raven

Defects filed against this service by stellar-raven, its largest agent consumer, from that project's own evaluation battery. These carry more signal than our internal detectors because the answer key is not ours.

Filed against usiEach is a defect their evaluation battery reproduced against a live surface of ours, with the probe and evidence recorded in their repository.9Across the project's lifetime
Their answer keyiStatus as THEIR records state it. A record stays reported-upstream until their re-verification runs; we never mark their findings closed.7 open1 declined by them
Fix shipped (our side)iScored by our own linked GitHub issue state. This is NOT the consumer's verdict and must never be rendered as one. A record can read reported-upstream on their side while our linked issue is closed: we shipped the fix and their re-verification has not run yet. Neither status is allowed to speak for the other.7/8↑ higher is better · Of the findings a fix applies to

Their answer key, by status

verified by them0fixed, awaiting their re-check0open7declined by them1

A record can read reported-upstream on their side while our linked issue is closed: we shipped the fix and their re-verification has not run yet. Neither status is allowed to speak for the other.

Defect flow: detector to surface to outcome

Every finding in the ledger traced through the system: which detector caught it, which surface it lives on, and whether it closed. Ribbon thickness is the count; whole-ledger, not a sample.

Findings traced from detector, through surface, to outcomeretrieval → Cleared: 335engine-a-recall → retrieval: 284directory → Cleared: 89engine-d-demand → directory: 61engine-d-demand → retrieval: 47nightly-completeness → directory: 28contract → Cleared: 27raven-routing → consumer: 24nightly-drift → contract: 23consumer → Cleared: 19nightly-battery → corpus: 18corpus → Cleared: 18scf-crosscheck → scf: 16scf → Cleared: 16code → Cleared: 10engine-e-contract → contract: 8contract → Verified: 7nightly-note-freshness → code: 6consumer → Open: 5golden-eval → retrieval: 4nightly-field-population → code: 4nightly-claims → contract: 3raven-loop → code: 1code → Open: 1engine-a-recall: 284 findingsengine-a-recall 284engine-d-demand: 108 findingsengine-d-demand 108nightly-completeness: 28 findingsnightly-completeness 28raven-routing: 24 findingsraven-routing 24nightly-drift: 23 findingsnightly-drift 23nightly-battery: 18 findingsnightly-battery 18scf-crosscheck: 16 findingsscf-crosscheck 16engine-e-contract: 8 findingsengine-e-contract 8nightly-note-freshness: 6 findingsnightly-note-freshness 6golden-eval: 4 findingsgolden-eval 4nightly-field-population: 4 findingsnightly-field-population 4nightly-claims: 3 findingsnightly-claims 3raven-loop: 1 findingsraven-loop 1retrieval: 335 findingsretrieval 335directory: 89 findingsdirectory 89contract: 34 findingscontract 34consumer: 24 findingsconsumer 24corpus: 18 findingscorpus 18scf: 16 findingsscf 16code: 11 findingscode 11Cleared: 514 findingsCleared 514Open: 6 findingsOpen 6Verified: 7 findingsVerified 7

Read left to right: a detector produces findings, they land on a surface, and they end Cleared, Open, or Verified. A fat ribbon into Open is a surface carrying real debt; a fat ribbon into Cleared is a detector whose class has been closed.

Where open recall misses die

Each open recall finding replayed live and classified at the FIRST stage that fails - mutually exclusive classes with different owners, not a funnel or a sequence.

0 of 0 open recall findings replayed

passing0 · 0%no longer reproduces · owner: ledger
ranking0 · 0%returned, but below top-3 · owner: ranking
admission0 · 0%not returned for the query at all · owner: admission / matching
identity0 · 0%own exact name does not return it · owner: indexing / identity
corpus0 · 0%not in the directory at all · owner: coverage / curation

Row quality: the evidence behind each record

Every project row scores on five facts we either hold or don't: a provenance basis, a date, a source URL, a type, and a link. A low score names exactly what is missing.

Mean row evidence scoreiPer row: the count of five BINARY evidence facts present, times 20 - a strong provenance basis, a status date, a source URL, at least one type, at least one link. Scores land only on 0/20/40/60/80/100. 100 means all five present; it does NOT rate the project, only how well we can back what we publish about it.87%↑ higher is better · across all 1100 rows, a census, not a sample

Status provenance, strongest evidence first

site-liveness506human-verified155repo-activity115source-inherited100product-integration58onchain-activity39unverified8operator-announcement3

Deployment fact (sls-079): 903 unknown · 79 mainnet · 2 testnetiWhich network a product is deployed on, as a separate fact from lifecycle status. Populated ONLY from evidence (verified mainnet contract joins, on-chain readings, human-verified operator artifacts); unknown is the honest default and a work queue, never a score. The gap matrix carries the prominent rows to work first.

Most rows rest on site-liveness - a page answered. That is the weakest honest basis we serve, and moving rows up this ramp is the standing data job.

What is missing, across all rows

strongBasis
617
dated
3
sourced
15
typed
2
linked
10

Every row as one dot. The marked region is the curation queue.

Prominence vs evidence, one dot per directory row0/51/52/53/54/55/50255075100prominent + weak evidence: 1 rows · work these first
evidence facts held (of 5) ↑curated prominence →

The code index: what it holds, how deeply we know it

2,920 curated repos (claimed by a project or a tracked builder) plus a 10,018-row Electric Capital tail indexed for completeness. The charts read over the curated index only - mixing the tail in made the curated index look unscanned when it is not.

Repo score distribution, curated indexirepoScore (0-100) grades freshness, traction and builder authority. A long low tail is EXPECTED in an open ecosystem - hackathon one-offs and early experiments are real code references worth indexing; the score is what keeps them ranked below production repos.

0–990–99

Commit activity, curated indexiDerived from each repo's last commit: active (<=90d), slowing (<=1y), dormant (older), archived, or unknown (no commit date held - not knowing is its own state, never counted as dormant).

dormant1,256slowing936active795unknown90archived34

Languages

TypeScript
1,106
unknown
450
Rust
448
JavaScript
350
Go
109
Python
105

How deeply we know each layeriThree different jobs with three different denominators. Depth scanning aims at the whole curated index. Knowledge notes are a hand-curated research layer being built over the highest-scored repos - a small number is early progress, not missing homework. Mainnet joins are deliberately strict: only a verified on-chain attribution counts, so the number grows slowly and every unit of it is proof.

Depth-scanned2,912 / 3,111 · 94%

Curated repos with a code-depth reading (entry files, symbols, SDK usage). The scan waves aim at all of them.

Deep-researched notes151 / 408 · 37%

Hand-curated dated facts with sources, written over the top-scored pool one repo at a time. An enrichment layer under construction, newest additions first.

Verified mainnet joins107 / 5,660 · 2%

Deployable contracts with a PROVEN on-chain attribution. Strict by design: absence is absence of a join, never proof of disuse.

Notes pool, honestly split: of 408 pool repos, 151 carry dated facts, 195 were examined and yielded nothing durable (judged, recorded internally), and 62 are still unexamined. A judged repo is not a gap.

Plus the Electric Capital tail: 7,935 of 10,018 rows scanned opportunistically as budget allows - indexed for completeness, no coverage target attached.

Highest-graded repos

Lessons and research

Every recurring defect class was written up when it was found, and every human-verified correction carries a committed receipt. These are the documents behind the numbers above.

lessonsauditsreceiptscumulative total (41)

Correction receiptsiA human-verified status change commits its evidence: the URL fetched, the time, response identity headers, and the exact markers on the page that decided the verdict. Re-run the capture to diff what a page says now against what it said then.

Audits

Trends

Daily history appended by the eval pipeline and committed, red days included. Battery probe counts rotate with the daily banks, so the pass line moves by design; the failure line and the ratchets are the signal.

battery passbattery failbattery errorsopen maps ratchet (line, lower is better)

SCF funding cross-check

No project overstates or understates SCF membership, at the project level OR the round level, against the fund's own directory.

measured 2026-08-28 · 8d ago · baseline
0
bad membership claims (project + round level)
at target
  • 310 matched records checked against 500 SCF projects
  • 0 overstated / 0 understated at project level
  • 0 round-level overclaims across 309 verified claims

Contract honesty probe (Engine E)

Documented params do something, undocumented values are rejected. The contract a stranger hits behaves as written.

measured 2026-08-28 · 8d ago · baseline
0
violations across 807 params + fields probed on 37 operations
at target
  • 0 param(s) documented but silently ignored
  • 0 param(s) accepting values the spec forbids
  • spec 1.9.1 at measurement; re-run on every deploy

Curated canonical repos resolve

Every repo we call authoritative is indexed and carries code signals — a curated name that matches no row silently degrades the query it was written for.

measured 2026-08-31 · 5d ago · baseline
44/44
44/44 curated canonical repos indexed and code-scanned
at target
  • 0 absent (curated name matches no row)
  • 0 indexed but no code signals — invisible to code-evidence ranking and to the tier gate

Computed values reach a serving path

Every field our machinery computes is read by something that shapes an agent's answer — a value nothing consumes cannot change what anyone is told.

measured 2026-08-31 · 5d ago · on-deploy
9/9
9/9 computed fields reach a serving path, or an engine whose output is served
at target
  • a script or a test does not count as consumption — that is how codeProofTier passed for months while only a report called it

Guard lanes that actually run

Every automated lane in the repo has completed a real run recently — a guard that never executes is a promise, not a check.

measured 2026-09-04 · 1d ago · weekly
84/86
84/86 automated lanes have a green run inside their own cadence
below target
  • FAILING: api-drift.yml — 2 failed run(s) since its last green 2d ago — it is trying and losing, not idle; dies at "Field-population guard (values arrive, not just shape)"
  • FAILING: workflow-health.yml — 5 failed run(s) since its last green 1d ago — it is trying and losing, not idle; dies at "Mirror drift (scout-mcp)"
  • judged against the workflow file as it stands: runs from a since-edited or never-merged version are ignored, so a fixed lane stops being red
  • a lane that exits 1 to report a finding is the guard working, and is not counted here

Type errors in scripts/

The scripts that write to the production database are type-checked, and the backlog of known errors only shrinks.

measured 2026-08-31 · 5d ago · on-deploy
63
63 type errors in scripts/, 63 of them baselined — new ones fail the build
at target
  • identity is file + error code + message, without line numbers, so an unrelated edit above an error does not read as a regression
  • a count of 0 here means the ratchet is finished and the guard can become a plain tsc gate

Served values a consumer can date

Every value we serve carries a date that covers IT — or says plainly that it cannot be dated. A value wearing a neighbour's timestamp is a wrong answer with a citation.

measured 2026-08-31 · 5d ago · on-deploy
46/80
46/80 served values are dated in their own scope, or documented as undatable
at target
  • explainRepo.answerAsOf is the pattern: NULL for a DeepWiki answer, because DeepWiki exposes no index date and inventing one would make an unknown look measured
  • scoping cuts both ways: verifyClaim's confidence.ageDays dates the numbers inside confidence but NOT the root-level verdict beside it — so verifyClaim.verdict is carried as named debt, not excused; an admission only counts when the description speaks to datability itself ('no index date'), never bare 'null' or 'unknown'
  • limit: a scope holding one date is treated as dating every value in it, so this catches the sharper shape only — a value with NO date in scope while other objects in the response carry dates

SCF-funded projects served

Every SCF-funded project (round-badged, so provably funded) is in the directory an agent searches.

measured 2026-09-01 · 4d ago · weekly
500/500
500/500 SCF projects served — 0 absent, all carrying a funding-round badge
at target
  • 0 SCF-funded projects the directory does not serve
  • DefiLlama: 0 missing of 42 Stellar-listed
  • absent entries are human-reviewed SEEDS, never bulk-created

Code-depth calibration

Repo depth grades agree with independent code analysis where both exist — on a sample large enough for the rate to mean something.

measured 2026-07-10 · 57d ago · baseline
100% (n=28)
agreement on 28 independently graded answer-key repos of 86 (deepwiki 19 + grok 9)
at target
  • 0 disagreements
  • two graders, recorded per row and never pooled silently: DeepWiki for the few answer-key repos it has indexed, Grok (agentic, reads the repo) for the rest
  • the binding constraints are EXTERNAL: DeepWiki's index covers ~4 of the 86 long-tail answer-key repos, and the Grok balance exhausted mid-run (HTTP 402) after 9 verdicts — topping it up or DeepWiki indexing more of the key is what grows n
  • sample meets the 20-repo floor

Consumer interlock (Raven)

The #1 consumer's discovery index tracks our contract, checked from OUR side too, with grace for their re-baseline cadence.

measured 2026-08-28 · 8d ago · on-deploy
29/33
operations discoverable in the consumer catalog
at target
  • 4 op(s) lagging within the 10-day re-baseline grace window (expected)
  • 0 op(s) missing beyond grace
  • contract 1.9.1 at measurement

Recall floors (Engine A)

Generated known-item probes per bucket stay above their red-line floors, recall can't silently erode.

measured 2026-09-01 · 4d ago · weekly
8/8
buckets at/above floor · 2,235 probes
at target
  • all buckets above floor

Real-demand OK-rate (Engine D)

The queries real consumers actually sent are replayed live and keep answering at or above the committed floor. A miss on real demand outranks any synthetic finding.

measured 2026-09-01 · 4d ago · weekly
99%
of the top 250 of 4,072 distinct real queries · floor 80%
at target
  • 9,254 real-consumer calls in the 14-day window
  • 2 queries missing today, the standing fix queue
  • floor 80% is a ratchet from a prior reading, it may only move up

Golden retrieval eval

Known-true questions (answer key derived from the canonical directory) keep passing after every ship.

measured 2026-09-04 · 1d ago · on-deploy
51/51
golden questions passing (scored)
at target
  • full pass
  • 2 N/A by design (live-source questions the static corpus doesn't answer)
  • re-run on every production deploy + weekly

Corpus hygiene (S5-S8)

The research corpus stays clean: junk URLs, broken titles, staleness and mirror drift are swept weekly, and a sweep that returns no reading counts as unknown, never as clean.

measured 2026-09-01 · 4d ago · weekly
5/5
sweeps clean · 10,132 chunks / 2,070 docs
at target
  • s7 coverage: 2 source(s) structurally undateable (declared)
  • all sweeps clean

Consumer path (through Raven)

The canonical questions answer correctly through the REAL Raven gateway: routing, our op, the response envelope and coaching, not just our direct API.

measured 2026-08-28 · 8d ago · on-deploy
47/47
golden Qs answered through the live gateway · 100%
at target
  • every gradeable golden question answers correctly through Raven (47 checked)

Improvement ledger

Every quality detector's findings land in one tracked backlog by surface. A backlog is fine; a HIGH-severity finding neglected past 30 days is the failure, that's this row's red line.

measured 2026-09-04 · 1d ago · weekly
5 open
0 high · 0 in a wave · 521 closed, 223 on evidence
at target
  • open by surface: consumer 5 · retrieval 0 · code 0 · directory 0 · scf 0 · contract 0 · corpus 0
  • 527 tracked · 0 in-wave · 7 verified · 514 auto-cleared (detector stopped flagging)
  • closure basis: 7 verified deliberately · 216 cleared on a live re-probe that PASSED · 298 cleared only because a detector stopped reporting
  • that last group is the re-probe backlog, not a result: a detector going quiet is indistinguishable from a gap nobody asked about again. Spot-check 2026-08-31 — kutana, etesia and octopos sit in it, are still absent from the directory, and each carries SCF round badges.
  • no high-severity finding neglected past 30 days
  • all detectors reporting within 10d, every open finding is a confirmed one

Response-shape opacity (build-enforced)

No silent opacity: every response object declares its shape. Grandfathered open maps are counted here and the count may only fall.

measured 2026-09-04 · 1d ago · on-deploy
0
grandfathered open maps (additionalProperties: true) still undeclared
at target
  • 0 response objects still return an undeclared shape
  • the ratchet fails the build if that number rises; it may only fall
  • a NEW operation shipping without a declared shape fails the build outright
  • scripts/contract/check-schema-opacity.ts, wired into contract:check

Match labelling (build-enforced)

Every q-taking operation declares HOW it matched, so a caller can tell an exact hit from a semantic neighbour. Enforced in CI.

measured 2026-09-04 · 1d ago · on-deploy
0
q-taking operations still missing a match label, of 4 tracked
at target
  • 4 q-taking operation(s) tracked, 0 unlabelled
  • a new q-taking operation shipping without a match label fails the build
  • scripts/contract/check-honesty-layer.ts, wired into contract:check

How this page works: scheduled runs measure, write their JSON evidence to improvements/ and commit it; this page statically renders those committed files, so a number here changes only when a new dated artifact lands. Entity counts are a census of the collections; guard rows are point-in-time measurements that carry their own age and go stale rather than silently reading as current. The consumer-side interlock conventions (spec-as-discovery-index, version handshake, cadence contract) are specified in docs/interlock-spec.md.