Measured, not asserted

Data quality

Every number here is read from a committed engine artifact, so this page cannot say anything the runs did not measure. Each figure links its evidence and carries the date it was taken. The standing promises live in DATA_SLA.md.

Verdict: where this data stands right now

Derived from the guard states below, not written by hand.

At target
12
Below target
4
Needs re-measure
3
Open findings
8
Waiting on upstream
17

At target = passing on fresh evidence. Below target = measured, with an open work queue. Needs re-measure = evidence older than its own window. Open findings are ours and still reproducing; waiting on upstream is carried, not ours to fix.

11 of those have been waiting longer than 20 days — the oldest for 71. A re-baseline that has not happened in 71 days is not a wait any more.

Safe to rely on

value · days since measured

Row coverage: this page reads all 1133 project rows from the unranked listing, a census, so no row hides by being hard to retrieve.

Progress against the quality plan

Phase status is read from QUALITY.md itself, a phase cannot show green here without being green there. Remaining work is shown in the same weight as completed work.

3/6 phases
P0P1P2P3P4P5
P0Name the classes, lock the first invariantDone
evidence ›
Evidence: check-schema-opacity.ts (the 47 baselined open maps were paid down to ZERO in #1092; the ratchet now holds the floor at 0), QUALITY.md itself.
P1Honesty layer, eval integrity, the dashboardDone
evidence ›
Evidence: specs/honesty-baseline.json (debt 0), scripts/eval/eval-baselines.json, improvements/quality/*.json.
P2Entity truth: issuers, receipts, enumerations, dedupeDone
evidence ›
Evidence: battery slice G (enumeration integrity), slice H (verify grades itself), improvements/receipts/, improvements/quality/entities.json (repos.duplicateRows = 0), scripts/data/check-repo-dupes.ts in enrich-repos.yml + dedupe-repos.yml.
P3Earned autonomyIn progressevery lane that writes to production has earned the intervention-free threshold — 10 → 76 lanes · 20% of the way from 0Remaining: five lanes at Stage 2 (enrich-tvl, scan-repo-code, refresh-research-corpus on 2026-09-13; enrich-repo-activity and refresh-stablecoins on 2026-09-14, each on an end-state claim proven twice in its own log), 7 eligible at 4+ clean weeks. Detection is autonomous; …
evidence ›
Remaining (full): five lanes at Stage 2 (enrich-tvl, scan-repo-code, refresh-research-corpus on 2026-09-13; enrich-repo-activity and refresh-stablecoins on 2026-09-14, each on an end-state claim proven twice in its own log), 7 eligible at 4+ clean weeks. Detection is autonomous; repair has a lane (repair-lane.yml, one open ledger row per day) and the four hand-run tracks have one (task-lane.yml: notes Tue, claims Wed, packets Thu, gap Fri) — both PR-only, headless agent, no production secrets, protected or allowed paths enforced by a diff-guard, the attempt log on main as the end-state claim. First-day proof: the repair lane found and fixed a real defect (lulpay's SCF record, #1568); the task lane delivered a 40-row packet (#1574); PR creation needs the repo to allow Actions to open PRs. Stage 3 stays closed until the lanes have weeks.Shipped so far: the daily pipeline now rebuilds its own quality artifacts and commits them, and the stale-finding sweep re-probes the ledger instead of letting counts drift. 2026-08-28: the FIRST bounded agent lane is live — the deployment-evidence gap (sls-079) is worked mechanically by the weekly curation pass via the operator-toml chain (the project's own stellar.toml -> declared code+issuer -> confirmed on Horizon mainnet; full chain or abstain, basis labeled "operator-toml" so a machine stamp never impersonates a human one). The lane reproduces the 2026-08-28 hand-worked queue's mechanical half; judgment cases (operator docs, bundles) stay human. 2026-08-29: the closure rule's METRIC exists — repeat-class rate is computed from the ledger and published on /quality (first measurement: 30-day rate 100%, 168/168 new findings in already-seen classes; lifetime 98.7% across 6 classes + meta-eval. The treadmill, now with a number).
P4Basis strength at scaleIn progressweak bases under 50% of Live rows — 54.7 → 50 % of Live rows on a weak basis · 87% of the way from 86Remaining: done bar is weak bases under 50% of Live rows; 453 of 830 (54.6%) on 2026-09-13 — read strongBasisSplit for today's figure. The machine levers are measured out (onchain-eligible 5, repo-activity 0); the ~39 rows that must move are the owner triage table (relink / Inactive / leave), human work.
evidence ›
Shipped so far: 2026-08-31→09-01 — the basis-upgrade lanes exist and have run against the population: evidence A (asset movement deltas between two dated stellar.expert readings), B (DeFiLlama TVL ≤14d) and C (Horizon, same day — issuer payments asset-matched over 20 records, XLM-pair trade fallback), every probe trinary (hit / checked-empty / could-not-check reaches the summary and the exit code). 35 rows now rest on onchain-activity; a two-auditor pass (agent + Grok) reverted 4 uncorrected-probe upgrades, one of which re-earned its upgrade the same run under the corrected rule. Asset keys joined from the stablecoin registry (12) and operators' own stellar.toml (7) so deltas compound weekly. Death receipts: 42 stamped, 8 retracted after audit (a live-200 page cannot stand as "observed dead"), 2 re-stamped once their domains went hard-dead. Untyped 59→2 (both honest residuals). Weak share: read strongBasisSplit and strongByBasis in improvements/quality/entities.json — the number is no longer hand-printed here, because on 2026-09-05 this file said 84%, the board 62.7% and knownLimitations 62% for one SLO. The 2026-09-04 drop (794→617 weak) was 173 rows on two NEW evidence tiers (repo-activity, product-integration) plus 4 within the pre-existing onchain-activity tier: real evidence, and a change in what counts — reported as two numbers from now on, never as the ratchet falling.
P5The knowledge layer consumers keep asking forIn progressthe curated-pool note floor, examined or noted — 371 → 373 curated-pool repos · 99% of the way from 0Remaining: every curated-pool row (repoScore ≥ 50) carries a public note or a triage verdict as of 2026-09-14 (waves 4–6: 345 → 394 public-noted repos on the detector; the DB floor follows the daily backfill). Still open: contracts as first-class joined entities (11 of the 308 expected-tier repos).
evidence ›
Shipped so far: 2026-09-01 — the curated knowledgeNotes registry grew 16→~29 repos, every note dated and source-cited, including the supersession facts consumers actually ask for as prose (stellar/go → go-stellar-sdk, the Horizon monorepo split, the js-sdk deprecation chain, protocol ceilings). explainRepo now answers from a dated note ahead of an undated DeepWiki walkthrough (answerSource: "knowledge-note", answerAsOf RFC 3339), matched by exact identifier or by hand-authored trigger phrases; the matcher was hijack-hardened (citation URLs and bare domains can no longer route a note). sls-080 closed on the consumer's own probe and independently re-verified by Raven on 2026-09-01. 2026-09-04/05 — the product-level knowledge sls-023 asked for (which product is actually issued on Stellar, by whom, with what controls) exists as a verified RWA registry: 97 tokens re-verified from the issuer's own stellar.toml or the Soroban contract itself, feeding products, deployment and on-chain controls on project rows, pinned hourly on production (#1298–#1306). 2026-09-05 — supersededBy / deprecatedAt / supersessionKind are FIELDS on repo rows: 50 notes carried the prose, 34 became a curated dated map keyed by the superseded repo, successorRepo is derived from it, and a test holds prose and fields together (spec 1.9.36, #1307).

Reading this as an agent?

Every number on this page is served as JSON: the verdict block first, then the north star with its age, per-operation contract state, known limitations, the gap matrix with real identifiers, the miss funnel, consumer findings, guard state and the trend history. No parameters, no key. Cached one hour and served stale up to a day while revalidating, so read meta.measuredAt, not the clock.

GET /api/quality

North star: full-surface audit ok-rate

Hundreds of cold, natural probes across every retrieval surface, graded against ground truth. The one number the whole engine system optimizes.

Reading note: the 9 points were measured over different probe counts (597, 648, 198, 2164, 2235, 2229, 2247), so they do not form a comparable trend line

Known limitations: read before relying on this data

Derived from the measurements below, not written by hand: if a number improves, the entry changes or disappears. Each says what to do instead.

project status459 of 1004 rows (census)

46% of lifecycle statuses rest on the weakest honest bases: a page answered (site-liveness) or a value inherited from a source (source-inherited).

Instead: Weigh statusBasis and statusAsOf on every row; treat human-verified and onchain-activity as the strong tiers, and verify a Live claim against the row's statusSourceUrl before repeating it.

project types6 of 1004 sampled rows untyped

Some rows carry no type, so an exact ?type= enumeration cannot see them even when the project belongs to that vertical.

Instead: For discovery, combine ?type= with a q= search; an empty typed result is a statement about our tagging, not about the ecosystem.

repo knowledge notes342 of 373 rows in the curation pool (curated index, repoScore >= 50)

Curated, dated repo facts (knowledgeNotes) exist on a minority of indexed repositories.

Instead: Absence of notes is absence of curation, never evidence about the repo; fall back to codeVerified and activity fields.

Open findings by surface. A work queue, not an outage

code
8
retrieval
0
scf
0
contract
0
consumer
0
directory
0
corpus
0

An agent can fetch all of this, limitations, surface health, guard state, the score definitions and the trend, from GET /api/quality, so trust calibration does not require reading a webpage.

Gap matrix: what is missing, by entity and field

One row per hole. Counts are samples with their denominator, and every row carries real identifiers so the gap can be worked or independently checked.

repo mainnet contract join359 / 378 missing (95%)

Without a verified contract join we can say code exists, never that it is USED on mainnet. Denominator is the expected tier only — 5380 further low-signal deployable repos (demos/tutorials/experiments) are deliberately excluded: absence of a mainnet join there is expected-normal, and sweeping them in overstated this gap ~10×. Closed by: Contract attribution during the scan wave. Absence is absence of a join, never proof of disuse.

e.g. fazzatti/colibri, stellar/stellar-horizon, Shamba-Records-Limited/microvault, hyperledger-solang/solang, Phoenix-Protocol-Group/phoenix-contracts, stellar/sep45-reference

repo knowledgeNotes12378 / 13437 missing (92%)

Curated dated facts are what let an agent cite a repo claim. Without them only raw scan fields are available. Closed by: The repo-intel enrich pass, which writes dated notes with sources.

e.g. luanlabs/fluxity-interface, Stellar-Light/stellar-pay, SoroWill/sorowill-sdk, Certora/CertoraProver, Vellar-Wallet/vellar-sdk, ravendevteam/alternet

project deployment evidence915 / 1004 missing (91%)

sls-079: 'Live' says a product operates for users somewhere; it does NOT say which network it is deployed on. Of the 915 unknown rows, 201 are on-chain product types where the question applies and an agent will ask it; the other 714 are SDKs, wallets, security and analytics rows where unknown is the honest answer, not a gap. Closed by: Evidence only: a verified mainnet contract join, an on-chain activity reading, or a human-verified operator artifact (DEPLOYMENT_VERIFIED).

e.g. wisdomtree, redswan, spacewalk

project strongBasis480 / 1004 missing (48%)

site-liveness means only that a page answered. source-inherited means the value came from elsewhere. Neither is observation of the product. Of the weak rows, 6 hold an on-chain footprint (an issued asset, a joined contract, or a known deployment) and can earn onchain-activity from evidence; the other 474 are app-only and can only move through human verification — the ceiling of this row is people, not lanes. Closed by: Human verification with a receipt, an on-chain activity reading, or an operator announcement.

e.g. yellow-card, scorechain, solar-wallet, infstones

repo code depth reading2303 / 13437 missing (17%)

No depth reading means the repo was never scanned for real implementation signal. Closed by: A scan wave pass over the unscanned tail.

e.g. Dione-b/stellar-start, thiagocbalducci/neon-wallet-desktop, adletgamer/pakta, FernandoMay/stellar-energy-assets, karagozemin/stellar-drips, ManuelJG1999/Thalos

project typed6 / 1004 missing (1%)

An untyped row is invisible to exact ?type= enumeration and to the gaps axis, even when it belongs to that vertical. Closed by: Add the type via the curation TYPE_ADD pass, with the row's own description as evidence.

e.g. tilt-pay, stellar-vrf, rail402, pigfi, lendwise, lendoor

project sourced0 / 1004 missing (0%)

A lifecycle claim with no source cannot be re-checked by a caller. It is our assertion, not evidence. Closed by: Curate a dated source URL, or downgrade the basis to match the evidence we actually have.

Agent lanes: autonomy earned, not assumed

A lane is a bounded, evidence-only job an agent runs on schedule. Advancement is measured — a lane earns auto-merge only after consecutive weeks where a human reviewed and changed nothing.

Lane: operator-toml
8 stamps
deployment facts, full-chain-or-abstain
Intervention-free weeks
2 / 2
to Stage 2 (auto-merge for bounded work)
Last run
success
schedule · 2026-09-29
Corrections
0
human had to fix a stamp — resets the counter

Weeks are counted only from successful scheduled runs, and the counter is derived daily from the lane's live write-set diffed against the committed snapshot — a quiet failure reads as a red week, never a clean one. A stamp a human upgrades to human-verified stays clean; a stamp a human removes or changes resets the count to zero.

Lane autonomy — intervention-free weeks

Every workflow in this repo that can write to production data, and the weeks each has earned toward running its execute unattended. Built by reading the workflow files, counted from GitHub's own run history, and reset by the intervention log — no lane's number is asserted here.

A week counts only when the lane executed and nothing it wrote was corrected.

Lanes that write to production
76
derived from .github/workflows each run, not from a list
At 4+ clean weeks
10
eligible — the promotion is still a human call
Promoted to Stage 2
5
recorded in improvements/lanes/lanes.json with the lane's own end-state claim
Could not check
11
API refused; not counted as clean or broken
LaneCadenceRuns 8wWeeksLast correction
check-linkscron 23 2 * * *31 / 118—
enrich-onchaincron 0 6 * * 28 / 158—
enrich-tvlcron 0 5 * * 19 / 18—
refresh-research-corpuscron 0 6 * * *48 / 228—
scan-repo-codecron 17 */2 * * * · cron 43 3 1 * *298 / 1578—
sync-lumenloopcron 0 7 * * *51 / 28—
aggregate-feedbackcron 40 3 * * *45 / 27—
backfill-triage-tagscron 0 4 * * *43 / 27—
enrich-builderscron 30 6 * * *42 / 36—
enrich-repo-activitycron 20 4 * * *41 / 46—
enrich-reposcron 0 8 * * 17 / 786—
refresh-stablecoinscron 25 */6 * * *134 / 496—
enrich-builder-reposcron 0 5 * * 25 / 75—
raven-eval-paritycron 30 7 * * *24 / 195—
upgrade-status-basiscron 23 4 * * *29 / 145—
backfill-knowledge-notescron 23 3 * * *26 / 3432026-09-05 · The nightly step was gated to workflow_dispatch, so the 3 scheduled runs since the cron was added on 2026-09-02 executed nothing (18 completed runs in total, most manual dispatches; corrected 2026-09-05 from the runs' job steps); the gate was removed so schedule runs write. Found by the lane-autonomy counter's first run.
generated-recallcron 0 7 * * 04 / 13—
refresh-rwacron 37 */6 * * *90 / 232026-09-05 · The sls-023 close-out on issue #494 was overclaimed — 11 of the finding's 61 rows were treated as closing the whole probe. Recorded as a violated lesson ("verify before advertise (#494, twice)"); the close-out itself has not yet been reframed upstream, so this stays open as a correction against the lane that now owns those rows.
verify-published-packagescron 40 5 * * 23 / 53—
award-reconcilecron 30 5 * * *11 / 142—
basis-from-deploymentcron 30 8 * * 22 / 72—
curate-projectscron 30 7 * * 23 / 17122026-09-05 · A July STATUS_FIX entry for blend (Live→Live, basis site-liveness) re-stamped site-liveness on every execute, erasing the onchain-activity the basis lane awards; 16 entries carry a weak basis. Fixed: a weak curated basis never overwrites a strong lane-earned one when the status is unchanged.
basis-from-onchaincron 0 9 * * 21 / 131—
basis-from-productcron 0 10 * * 21 / 812026-09-04 · The lane's writes were correct, its read-back was not: STRONG_BASES still listed only three tiers, so the 58 rows this lane awarded `product-integration` were counted as weak and the board read the completed work as failed.
basis-from-repocron 30 9 * * 21 / 412026-09-04 · Same defect, same PR: the 115 rows this lane awarded `repo-activity` were excluded from the strong-basis metric, so weakLiveRows read 790 when it was 617.
upgrade-basis-onchaincron 0 9 * * 21 / 121—
apply-content-passdispatch0 / 00—
apply-status-batchdispatch0 / 00—
apply-wallet-availabilitydispatch0 / 60—
award-importdispatch0 / 270—
award-publishdispatch0 / 70—
award-rounddispatch0 / 540—
award-test-setupdispatch0 / 310—
backfill-code-domainsdispatch0 / 50—
backfill-code-tierdispatch0 / 00—
backfill-contract-basisdispatch0 / 20—
backfill-lifecycledispatch0 / 00—
backfill-project-githubdispatch0 / 20—
backfill-provenancedispatch0 / 20—
backfill-version-statusdispatch0 / 20—
curate-partnersdispatch0 / 60—
db-spacedispatch0 / 00—
dedup-projectsdispatch0 / 302026-09-05 · Its 11 Draft hides were overwritten 30 minutes later by curate-projects' DUPE_MERGES step, which forces status=Inactive on every shadow (scripts/data/curate-projects.ts, 'if (dupe.status !== "Inactive")'). Two lanes own one field with different verdicts; the July fold (Inactive + canonicalSlug) won silently. Rows stay hidden from search and the board either way; the wrong status on duplicates is the open class in improvements/drafts/2026-09-05-inactive-site-liveness-triage.md.
dedupe-reposdispatch0 / 50—
embed-projectsdispatch0 / 00—
enrich-entitiesdispatch0 / 90—
enrich-partner-onchaindispatch0 / 00—
enrich-partnersdispatch0 / 30—
enrich-scfdispatch0 / 340—
fix-dev-docs-titlesdispatch0 / 10—
fix-research-titlesdispatch0 / 50—
fix-scf-roundsdispatch0 / 40—
fix-zenexdispatch0 / 00—
gone-reposcron 40 6 * * *0 / 10—
import-replit-stablecoin-historydispatch0 / 20—
ingest-dora-evalsdispatch0 / 00—
ingest-ec-taxonomydispatch0 / 90—
link-canonical-slugdispatch0 / 00—
link-partner-projectsdispatch0 / 00—
mark-defunctdispatch0 / 00—
mark-inactive-projectsdispatch0 / 20—
migrate-research-url-hostdispatch0 / 00—
patch-rozodispatch0 / 00—
prune-denied-reposdispatch0 / 20—
regrade-reposdispatch0 / 250—
repair-lanecron 47 8 * * *0 / 20—
seed-blog-postsdispatch0 / 30—
seed-i3-mockdispatch0 / 00—
seed-partnersdispatch0 / 00—
seed-summit-sp-winnersdispatch0 / 40—
set-project-statusdispatch0 / 00—
set-prominencedispatch0 / 00—
set-typesdispatch0 / 00—
tansu-anchordispatch0 / 120—
task-lanecron 17 9 * * 2 · cron 17 9 * * 3 · cron 17 9 * * 4 · cron 17 9 * * 51 / 00—
vector-indexdispatch0 / 20—
how a week is counted ›

Elapsed time earns nothing. The weeks must be consecutive and must run up to this one, so a lane nobody has run sits at zero however long it has been quiet, and four scattered good weeks are not four clean weeks. Only runs the lane started ITSELF count — a hand-dispatched execute is a person operating the lane, and is reported here rather than counted. The run counts are runs, not writes: which of them wrote is what the week count is proven from. Every counted execute is proven from that run's own job steps, never from today's copy of the workflow file: a step that was skipped moved nothing, whatever the file says now. A “+” marks a floor: GitHub does not expose a run's commands, so a step the author named and that actually ran could not be classified. Corrections live in improvements/lanes/interventions.json, appended by the same PR that makes the correction.

Findings: what the engines caught

Every detector writes here. This is the work queue, not a score.

Open
8
Ours, still reproducing
Waiting on upstream
17
raven-catalog-lag 11 · raven-scorer 6
Cleared
653
stopped reproducing
Verified closed
7
Re-probed after the fix

The three states are disjoint and sum to 741. Cleared is NOT confirmation the fix works; only verified means it was deliberately re-probed after a fix.

Caught in the world
64
The product working: a dead link, a stale note, a project that shut down
Our instrument
17
Measurement or serving broke: a failed golden question, a spec that lies

Both are still-open rows across all three buckets above (open, waiting on upstream, and the refresh queue), split by what each finding IS rather than whose turn it is. A world finding is repaired by curation and is evidence the detectors do their job; an instrument finding is ours. Lifetime: 268 world, 473 instrument. The mapping is one table per detector and failure mode in src/lib/improvement-ledger.ts; an unmapped pair counts as instrument, never as the product working.

Recurred after a silence-close (30d)
38.7%
86 of 222 new findings
Lifetime
27.9%
207 of 741 findings
Reopened after closing
12
0 of them had been VERIFIED

A recurrence is a NEW finding on a surface-and-failure-mode pair we had already closed on silence — the detector went quiet and nobody re-probed. It is the closure rule's real question: did we close without repairing, and did the same kind of failure come back? Steady state is that rate at zero. Reopened counts the exact-id version and is a lower bound, because re-clearing a reopened finding erases its stamp. These numbers are expected to start ugly; publishing them is the point.

For context, the repeat-class rate — a finding whose §0 class (identity, taxonomy coverage, contract completeness…) already had any prior finding — is 99.5% over 30 days across 7 classes. With classes that broad it cannot fall, so it is reported as context rather than steered by.

By failure mode and state

openclearedverifiedrecall-miss
broken-link
coverage-gap
demand-miss
completeness-residual
repo-gone
api-drift
battery-coverage-weak
scf-round-overclaim
fewermoreempty cell = zero

How long the open ones have been open

≤7d
8
8–30d
0
31–60d
0
>60d
0

A tall bar on the right is the treadmill this page exists to end: detection outrunning remediation. An empty right bar just after a stale-findings sweep means old entries were re-probed and cleared, not remediated.

Consumer findings from Raven

Defects filed against this service by stellar-raven, its largest agent consumer, from that project's own evaluation battery. These carry more signal than our internal detectors because the answer key is not ours.

Filed against usiEach is a defect their evaluation battery reproduced against a live surface of ours, with the probe and evidence recorded in their repository.9Across the project's lifetime
Their answer keyiStatus as THEIR records state it. A record stays reported-upstream until their re-verification runs; we never mark their findings closed.8 open1 declined by them
Fix shipped (our side)iScored by our own linked GitHub issue state. This is NOT the consumer's verdict and must never be rendered as one. A record can read reported-upstream on their side while our linked issue is closed: we shipped the fix and their re-verification has not run yet. Neither status is allowed to speak for the other.3/8↑ higher is better · Of the findings a fix applies to

Their answer key, by status

verified by them0fixed, awaiting their re-check0open8declined by them1

A record can read reported-upstream on their side while our linked issue is closed: we shipped the fix and their re-verification has not run yet. Neither status is allowed to speak for the other.

Defect flow: detector to surface to outcome

Every finding in the ledger traced through the system: which detector caught it, which surface it lives on, and whether it closed. Ribbon thickness is the count; whole-ledger, not a sample.

Findings traced from detector, through surface, to outcomeretrieval → Cleared: 342engine-a-recall → retrieval: 289directory → Cleared: 188link-health → directory: 113engine-d-demand → directory: 61engine-d-demand → retrieval: 49code → Open: 49nightly-note-freshness → code: 46raven-routing → consumer: 36code → Cleared: 34contract → Cleared: 30nightly-completeness → directory: 29gone-repos → code: 29nightly-drift → contract: 25nightly-battery → corpus: 24corpus → Cleared: 24consumer → Cleared: 19consumer → Open: 17scf-crosscheck → scf: 16scf → Cleared: 16directory → Open: 15engine-e-contract → contract: 8contract → Verified: 7golden-eval → retrieval: 4nightly-field-population → code: 4nightly-claims → contract: 4supersession-freshness → code: 3raven-loop → code: 1engine-a-recall: 289 findingsengine-a-recall 289link-health: 113 findingslink-health 113engine-d-demand: 110 findingsengine-d-demand 110nightly-note-freshness: 46 findingsnightly-note-freshness 46raven-routing: 36 findingsraven-routing 36nightly-completeness: 29 findingsnightly-completeness 29gone-repos: 29 findingsgone-repos 29nightly-drift: 25 findingsnightly-drift 25nightly-battery: 24 findingsnightly-battery 24scf-crosscheck: 16 findingsscf-crosscheck 16engine-e-contract: 8 findingsengine-e-contract 8golden-eval: 4 findingsgolden-eval 4nightly-field-population: 4 findingsnightly-field-population 4nightly-claims: 4 findingsnightly-claims 4supersession-freshness: 3 findingssupersession-freshness 3raven-loop: 1 findingsraven-loop 1retrieval: 342 findingsretrieval 342directory: 203 findingsdirectory 203code: 83 findingscode 83contract: 37 findingscontract 37corpus: 24 findingscorpus 24consumer: 36 findingsconsumer 36scf: 16 findingsscf 16Cleared: 653 findingsCleared 653Open: 81 findingsOpen 81Verified: 7 findingsVerified 7

Read left to right: a detector produces findings, they land on a surface, and they end Cleared, Open, or Verified. A fat ribbon into Open is a surface carrying real debt; a fat ribbon into Cleared is a detector whose class has been closed.

Where open recall misses die

Each open recall finding replayed live and classified at the FIRST stage that fails - mutually exclusive classes with different owners, not a funnel or a sequence.

0 of 0 open recall findings replayed

passing0 · 0%no longer reproduces · owner: ledger
ranking0 · 0%returned, but below top-3 · owner: ranking
admission0 · 0%not returned for the query at all · owner: admission / matching
identity0 · 0%own exact name does not return it · owner: indexing / identity
corpus0 · 0%not in the directory at all · owner: coverage / curation

Row quality: the evidence behind each record

Every project row scores on five facts we either hold or don't: a provenance basis, a date, a source URL, a type, and a link. A low score names exactly what is missing.

Mean row evidence scoreiPer row: the count of five BINARY evidence facts present, times 20 - a strong provenance basis, a status date, a source URL, at least one type, at least one link. Scores land only on 0/20/40/60/80/100. 100 means all five present; it does NOT rate the project, only how well we can back what we publish about it.89%↑ higher is better · across all 1133 rows, a census, not a sample

Status provenance, strongest evidence first

site-liveness382human-verified281repo-activity117source-inherited77product-integration67onchain-activity50operator-announcement15package-release9unverified6

Deployment fact (sls-079): 915 unknown · 87 mainnet · 2 testnetiWhich network a product is deployed on, as a separate fact from lifecycle status. Populated ONLY from evidence (verified mainnet contract joins, on-chain readings, human-verified operator artifacts); unknown is the honest default and a work queue, never a score. The gap matrix carries the prominent rows to work first.

Most rows rest on site-liveness - a page answered. That is the weakest honest basis we serve, and moving rows up this ramp is the standing data job.

What is missing, across all rows

strongBasis
480
dated
0
sourced
0
typed
6
linked
50

Every row as one dot. The marked region is the curation queue.

Prominence vs evidence, one dot per directory row0/51/52/53/54/55/50255075100prominent + weak evidence: 0 rows · work these first
evidence facts held (of 5) ↑curated prominence →

The code index: what it holds, how deeply we know it

2,920 curated repos (claimed by a project or a tracked builder) plus a 10,018-row Electric Capital tail indexed for completeness. The charts read over the curated index only - mixing the tail in made the curated index look unscanned when it is not.

Repo score distribution, curated indexirepoScore (0-100) grades freshness, traction and builder authority. A long low tail is EXPECTED in an open ecosystem - hackathon one-offs and early experiments are real code references worth indexing; the score is what keeps them ranked below production repos.

0–990–99

Commit activity, curated indexiDerived from each repo's last commit: active (<=90d), slowing (<=1y), dormant (older), archived, or unknown (no commit date held - not knowing is its own state, never counted as dormant).

dormant1,352slowing1,014active888unknown120archived45

Languages

TypeScript
1,200
unknown
521
Rust
495
JavaScript
381
Python
123
Go
122

How deeply we know each layeriThree different jobs with three different denominators. Depth scanning aims at the whole curated index. Knowledge notes are a hand-curated research layer being built over the highest-scored repos - a small number is early progress, not missing homework. Mainnet joins are deliberately strict: only a verified on-chain attribution counts, so the number grows slowly and every unit of it is proof.

Depth-scanned3,186 / 3,419 · 93%

Curated repos with a code-depth reading (entry files, symbols, SDK usage). The scan waves aim at all of them.

Deep-researched notes342 / 373 · 92%

Hand-curated dated facts with sources, written over the top-scored pool one repo at a time. An enrichment layer under construction, newest additions first.

Verified mainnet joins122 / 5,758 · 2%

Deployable contracts with a PROVEN on-chain attribution. Strict by design: absence is absence of a join, never proof of disuse.

Notes pool, honestly split: of 373 pool repos, 342 carry dated facts, 29 were examined and yielded nothing durable (judged, recorded internally), and 2 are still unexamined. A judged repo is not a gap.

Plus the Electric Capital tail: 7,948 of 10,018 rows scanned opportunistically as budget allows - indexed for completeness, no coverage target attached.

Highest-graded repos

Lessons and research

Every recurring defect class was written up when it was found, and every human-verified correction carries a committed receipt. These are the documents behind the numbers above.

lessonsauditsreceiptscumulative total (78)

Lesson write-upsiEach file records defects found on one day: what broke, the root cause, and the invariant or probe added so the class cannot silently return.

Guarded
14/19
Every Guard: line names a check that exists
Unguarded
5
Say Guard: none — written down, never became a check

A lesson counts only when it became a check — a test, a guard script or a workflow that goes red when the class returns. Each file names its guard with a Guard: line; check-lessons-guarded verifies the file exists and fails on any that is unguarded, on every PR.

Trends

Daily history appended by the eval pipeline and committed, red days included. Battery probe counts rotate with the daily banks, so the pass line moves by design; the failure line and the ratchets are the signal.

battery passbattery failbattery errorsopen maps ratchet (line, lower is better)

SCF funding cross-check

No project overstates or understates SCF membership, at the project level OR the round level, against the fund's own directory.

measured 2026-09-27 · 5d ago · weekly
0
bad membership claims (project + round level)
at target
  • 340 matched records checked against 531 SCF projects
  • 0 overstated / 0 understated at project level
  • 0 round-level overclaims across 328 verified claims

Contract honesty probe (Engine E)

Documented params do something, undocumented values are rejected. The contract a stranger hits behaves as written.

measured 2026-09-27 · 5d ago · weekly
0
violations across 962 params + fields probed on 38 operations
at target
  • 0 param(s) documented but silently ignored
  • 0 param(s) accepting values the spec forbids
  • spec 1.9.54 at measurement; probed on every deploy, evidence committed weekly by engine-c-health

Curated canonical repos resolve

Every repo we call authoritative is indexed and carries code signals — a curated name that matches no row silently degrades the query it was written for.

measured 2026-08-31 · 32d ago · baseline
44/44
44/44 curated canonical repos indexed and code-scanned
at target
  • 0 absent (curated name matches no row)
  • 0 indexed but no code signals — invisible to code-evidence ranking and to the tier gate

Human-verified stamps still hold

Every human-verified packet stamp is re-probed weekly against its own deciding URL; a contradiction is a finding, not a silent stale claim.

measured 2026-09-30 · 2d ago · weekly
75/77
75/77 packet stamps re-probed at their own sourceUrl and still supported by it
at target
  • 2 could-not-check — a 403, a timeout or a client-rendered shell is a page we did not read, never a contradiction
  • read-only: this guard files a finding for a human and never writes a status, because reading a marketing banner as a product state is the failure it exists to catch

Computed values reach a serving path

Every field our machinery computes is read by something that shapes an agent's answer — a value nothing consumes cannot change what anyone is told.

measured 2026-09-14 · 18d ago · on-deploy
9/9
9/9 computed fields reach a serving path, or an engine whose output is served
needs re-measure
  • a script or a test does not count as consumption — that is how codeProofTier passed for months while only a report called it

Guard lanes that actually run

Every automated lane in the repo has completed a real run recently — a guard that never executes is a promise, not a check.

measured 2026-10-02 · 0d ago · weekly
104/107
104/107 automated lanes have a green run inside their own cadence
below target
  • STANDING-SIGNAL: gone-repos.yml — 19 run(s) red at its declared signal step since its last green 17d ago — the lane works; what it measures has failed every run and nobody has acted (step "Probe")
  • STANDING-SIGNAL: raven-eval-parity.yml — 3 run(s) red at its declared signal step since its last green 1d ago — the lane works; what it measures has failed every run and nobody has acted (step "Truth battery (guard D — rotating probes, curated answer key)")
  • STANDING-SIGNAL: task-lane.yml — 7 run(s) red at its declared signal step since its last green 13d ago — the lane works; what it measures has failed every run and nobody has acted (step "Verify the end state")
  • judged against the workflow file as it stands: runs from a since-edited or never-merged version are ignored, so a fixed lane stops being red
  • a lane that exits 1 to report a finding is the guard working, and is not counted here

Type errors in scripts/

The scripts that write to the production database are type-checked, and the backlog of known errors only shrinks.

measured 2026-09-14 · 18d ago · on-deploy
29
29 type errors in scripts/, 29 of them baselined — new ones fail the build
needs re-measure
  • identity is file + error code + message, without line numbers, so an unrelated edit above an error does not read as a regression
  • a count of 0 here means the ratchet is finished and the guard can become a plain tsc gate

Served values a consumer can date

Every value we serve carries a date that covers IT — or says plainly that it cannot be dated. A value wearing a neighbour's timestamp is a wrong answer with a citation.

measured 2026-09-14 · 18d ago · on-deploy
49/83
49/83 served values are dated in their own scope, or documented as undatable
needs re-measure
  • explainRepo.answerAsOf is the pattern: NULL for a DeepWiki answer, because DeepWiki exposes no index date and inventing one would make an unknown look measured
  • scoping cuts both ways: verifyClaim's confidence.ageDays dates the numbers inside confidence but NOT the root-level verdict beside it — so verifyClaim.verdict is carried as named debt, not excused; an admission only counts when the description speaks to datability itself ('no index date'), never bare 'null' or 'unknown'
  • limit: a scope holding one date is treated as dating every value in it, so this catches the sharper shape only — a value with NO date in scope while other objects in the response carry dates

SCF-funded projects served

Every SCF-funded project is in the directory an agent searches, with funding read from SCF's own submission records.

measured 2026-10-01 · 1d ago · weekly
521/531
521/531 SCF projects served — 10 absent, 3 of them SCF awardees
below target
  • 10 SCF-funded projects the directory does not serve
  • DefiLlama: 1 missing of 42 Stellar-listed
  • absent entries are human-reviewed SEEDS, never bulk-created
  • absent: open-rwa-vault-infrastructure-nis (round 45)
  • absent: canary-protocol-y4l (round )
  • absent: pods-dxl (round )

Code-depth calibration

Repo depth grades agree with independent code analysis where both exist — on a sample large enough for the rate to mean something.

measured 2026-07-10 · 84d ago · baseline
100% (n=28)
agreement on 28 independently graded answer-key repos of 86 (deepwiki 19 + grok 9)
at target
  • 0 disagreements
  • two graders, recorded per row and never pooled silently: DeepWiki for the few answer-key repos it has indexed, Grok (agentic, reads the repo) for the rest
  • the binding constraints are EXTERNAL: DeepWiki's index covers ~4 of the 86 long-tail answer-key repos, and the Grok balance exhausted mid-run (HTTP 402) after 9 verdicts — topping it up or DeepWiki indexing more of the key is what grows n
  • sample meets the 20-repo floor

Consumer interlock (Raven)

The #1 consumer's discovery index tracks our contract, checked from OUR side too, with grace for their re-baseline cadence.

measured 2026-10-01 · 1d ago · weekly
27/33
operations discoverable in the consumer catalog
below target
  • 0 op(s) lagging within the 10-day re-baseline grace window (expected)
  • 0 op(s) missing beyond grace
  • 3 op(s) in the catalog that our own routing words do not reach — our fix
  • 3 op(s) callable but absent from the consumer's discovery index past the 10-day grace window (getQualityReport, getRwaAssets, verifyClaim) — an agent that does not already know the name cannot find them, and no wording of ours can change that
  • contract 1.9.54 at measurement

Recall floors (Engine A)

Generated known-item probes per bucket stay above their red-line floors, recall can't silently erode.

measured 2026-09-27 · 5d ago · weekly
8/8
buckets at/above floor · 2,275 probes
at target
  • all buckets above floor

Real-demand OK-rate (Engine D)

The queries real consumers actually sent are replayed live and keep answering at or above the committed floor. A miss on real demand outranks any synthetic finding.

measured 2026-09-27 · 5d ago · weekly
100%
of the top 190 of 7,388 distinct real queries · floor 80%
at target
  • 64,185 real-consumer calls in the 14-day window
  • 0 queries missing today, the standing fix queue
  • floor 80% is a ratchet from a prior reading, it may only move up

Golden retrieval eval

Known-true questions (answer key derived from the canonical directory) keep passing after every ship.

measured 2026-10-01 · 1d ago · on-deploy
51/51
golden questions passing (scored)
at target
  • full pass
  • 2 N/A by design (live-source questions the static corpus doesn't answer)
  • re-run on every production deploy + weekly

Corpus hygiene (S5-S8)

The research corpus stays clean: junk URLs, broken titles, staleness and mirror drift are swept weekly, and a sweep that returns no reading counts as unknown, never as clean.

measured 2026-09-27 · 5d ago · weekly
5/5
sweeps clean · 10,788 chunks / 2,185 docs
at target
  • s7 coverage: 2 source(s) structurally undateable (declared)
  • all sweeps clean

Consumer path (through Raven)

The canonical questions answer correctly through the REAL Raven gateway: routing, our op, the response envelope and coaching, not just our direct API.

measured 2026-09-27 · 5d ago · weekly
47/47
golden Qs answered through the live gateway · 100%
at target
  • every gradeable golden question answers correctly through Raven (47 checked)

Improvement ledger

Every quality detector's findings land in one tracked backlog by surface. A backlog is fine; a HIGH-severity finding neglected past 30 days is the failure, that's this row's red line.

measured 2026-10-01 · 1d ago · weekly
8 open
8 high · 0 in a wave · 660 closed, 223 on evidence
below target
  • open by surface: code 8 · retrieval 0 · directory 0 · scf 0 · contract 0 · consumer 0 · corpus 0
  • 741 tracked · 0 in-wave · 7 verified · 653 auto-cleared (detector stopped flagging)
  • closure basis: 7 verified deliberately · 216 cleared on a live re-probe that PASSED · 437 cleared only because a detector stopped reporting
  • that last group is the re-probe backlog, not a result: a detector going quiet is indistinguishable from a gap nobody asked about again. Spot-check 2026-08-31 — kutana, etesia and octopos sit in it, are still absent from the directory, and each carries SCF round badges.
  • no high-severity finding neglected past 30 days
  • ⚠ raven-routing (17d, 17 open), evidence past the 10d window; 17 open finding(s) unconfirmed

Response-shape opacity (build-enforced)

No silent opacity: every response object declares its shape. Grandfathered open maps are counted here and the count may only fall.

measured 2026-10-02 · 0d ago · on-deploy
0
grandfathered open maps (additionalProperties: true) still undeclared
at target
  • 0 response objects still return an undeclared shape
  • the ratchet fails the build if that number rises; it may only fall
  • a NEW operation shipping without a declared shape fails the build outright
  • scripts/contract/check-schema-opacity.ts, wired into contract:check

Match labelling (build-enforced)

Every q-taking operation declares HOW it matched, so a caller can tell an exact hit from a semantic neighbour. Enforced in CI.

measured 2026-10-02 · 0d ago · on-deploy
0
q-taking operations still missing a match label, of 4 tracked
at target
  • 4 q-taking operation(s) tracked, 0 unlabelled
  • a new q-taking operation shipping without a match label fails the build
  • scripts/contract/check-honesty-layer.ts, wired into contract:check

How this page works: scheduled runs measure, write their JSON evidence to improvements/ and commit it; this page statically renders those committed files, so a number here changes only when a new dated artifact lands. Entity counts are a census of the collections; guard rows are point-in-time measurements that carry their own age and go stale rather than silently reading as current. The consumer-side interlock conventions (spec-as-discovery-index, version handshake, cadence contract) are specified in docs/interlock-spec.md.