Measured, not asserted
Data quality
Every number below is read from a committed engine artifact, this page cannot say anything the runs didn't measure. Each stat links its reproducible evidence; each figure carries the date it was measured. The standing promises behind these guards live in DATA_SLA.md.
Verdict: where this data stands right now
Derived from the guard states below, not written by hand. A guard is at target only when its evidence is both passing and fresh; below-target rows carry their own work queue, and aged evidence counts as unmeasured, never as passing.
Safe to rely on
- SCF funding cross-check · 0 · measured 2026-08-28
- Contract honesty probe (Engine E) · 0 · measured 2026-08-28
- Curated canonical repos resolve · 44/44 · measured 2026-08-31
- Computed values reach a serving path · 9/9 · measured 2026-08-31
- Type errors in scripts/ · 63 · measured 2026-08-31
- Served values a consumer can date · 46/80 · measured 2026-08-31
- SCF-funded projects served · 500/500 · measured 2026-09-01
- Code-depth calibration · 100% (n=28) · measured 2026-07-10
- Consumer interlock (Raven) · 29/33 · measured 2026-08-28
- Recall floors (Engine A) · 8/8 · measured 2026-09-01
- Real-demand OK-rate (Engine D) · 99% · measured 2026-09-01
- Golden retrieval eval · 51/51 · measured 2026-09-04
- Corpus hygiene (S5-S8) · 5/5 · measured 2026-09-01
- Consumer path (through Raven) · 47/47 · measured 2026-08-28
- Improvement ledger · 5 open · measured 2026-09-04
- Response-shape opacity (build-enforced) · 0 · measured 2026-09-04
- Match labelling (build-enforced) · 0 · measured 2026-09-04
- Row coverage · this page reads all 1100 project rows from the unranked listing, a census, so no row hides by being hard to retrieve
Below target, being worked
- Guard lanes that actually run is below target: 84/86 lanes, measured 2026-09-04
Progress against the quality plan
Phase status is read from QUALITY.md itself, a phase cannot show green here without being green there. Remaining work is shown in the same weight as completed work.
Reading this as an agent?
Every number on this page is served as JSON: the verdict block first, then the north star with its age, per-operation contract state, known limitations, the gap matrix with real identifiers, the miss funnel, consumer findings, guard state and the trend history. No parameters, no key. Cached one hour and served stale up to a day while revalidating, so read meta.measuredAt, not the clock.
North star: full-surface audit ok-rate
Hundreds of cold, natural probes across every retrieval surface, graded against ground truth. The one number the whole engine system optimizes.
Reading note: the 5 points were measured over different probe counts (597, 648, 198, 2164), so they do not form a comparable trend line
Known limitations: read before relying on this data
Derived from the measurements below, not written by hand: if a number improves, the entry changes or disappears. Each says what to do instead.
62% of lifecycle statuses rest on the weakest honest bases: a page answered (site-liveness) or a value inherited from a source (source-inherited).
Instead: Weigh statusBasis and statusAsOf on every row; treat human-verified and onchain-activity as the strong tiers, and verify a Live claim against the row's statusSourceUrl before repeating it.
Some rows carry no type, so an exact ?type= enumeration cannot see them even when the project belongs to that vertical.
Instead: For discovery, combine ?type= with a q= search; an empty typed result is a statement about our tagging, not about the ecosystem.
Curated, dated repo facts (knowledgeNotes) exist on a minority of indexed repositories.
Instead: Absence of notes is absence of curation, never evidence about the repo; fall back to codeVerified and activity fields.
Open findings by surface. A work queue, not an outage
An agent can fetch all of this, limitations, surface health, guard state, the score definitions and the trend, from GET /api/quality, so trust calibration does not require reading a webpage.
Gap matrix: what is missing, by entity and field
One row per hole. Counts are samples with their denominator, and every row carries real identifiers so the gap can be worked or independently checked.
Without a verified contract join we can say code exists, never that it is USED on mainnet. Denominator is the expected tier only — 5265 further low-signal deployable repos (demos/tutorials/experiments) are deliberately excluded: absence of a mainnet join there is expected-normal, and sweeping them in overstated this gap ~10×. Closed by: Contract attribution during the scan wave. Absence is absence of a join, never proof of disuse.
e.g. Andy00L/x402-autopilot, Bosun-Josh121/clevercon, davidmaronio/StellarPay402, ritik4ever/lodestar, velikanghost/heekowave, 402md/agentcard
Curated dated facts are what let an agent cite a repo claim. Without them only raw scan fields are available. Closed by: The repo-intel enrich pass, which writes dated notes with sources.
e.g. Andy00L/x402-autopilot, velikanghost/heekowave, thewoodfish/AgentCompute, 402md/agentcard, ashfrancis/chickenz, catmcgee/stellar-poker-cosnarks
sls-079: 'Live' says a product operates for users somewhere; it does NOT say which network it is deployed on. Of the 903 unknown rows, 199 are on-chain product types where the question applies and an agent will ask it; the other 704 are SDKs, wallets, security and analytics rows where unknown is the honest answer, not a gap. Closed by: Evidence only: a verified mainnet contract join, an on-chain activity reading, or a human-verified operator artifact (DEPLOYMENT_VERIFIED).
e.g. wisdomtree, redswan, spacewalk
site-liveness means only that a page answered. source-inherited means the value came from elsewhere. Neither is observation of the product. Of the weak rows, 34 hold an on-chain footprint (an issued asset, a joined contract, or a known deployment) and can earn onchain-activity from evidence; the other 583 are app-only and can only move through human verification — the ceiling of this row is people, not lanes. Closed by: Human verification with a receipt, an on-chain activity reading, or an operator announcement.
e.g. kale, allbridge, ondo, quicknode, dia, axelar
No depth reading means the repo was never scanned for real implementation signal. Closed by: A scan wave pass over the unscanned tail.
e.g. leocagli/pontepay, leocagli/tralalero-contracts, KingFRANKHOOD/soroban-sql-sync, KB2410/ai-wallet-manager, KaruG1999/ocean_request, karagozemin/frontend-app
A lifecycle claim with no source cannot be re-checked by a caller. It is our assertion, not evidence. Closed by: Curate a dated source URL, or downgrade the basis to match the evidence we actually have.
e.g. brale, ylds, mydatacoin, transfermole, the-blue-marble, scam-flagging-system
An untyped row is invisible to exact ?type= enumeration and to the gaps axis, even when it belongs to that vertical. Closed by: Add the type via the curation TYPE_ADD pass, with the row's own description as evidence.
e.g. boundless-bounties, deb
Agent lanes: autonomy earned, not assumed
A lane is a bounded, evidence-only job an agent runs on schedule. Advancement is measured — a lane earns auto-merge only after consecutive weeks where a human reviewed and changed nothing.
Weeks are counted only from successful scheduled runs, and the counter is derived daily from the lane's live write-set diffed against the committed snapshot — a quiet failure reads as a red week, never a clean one. A stamp a human upgrades to human-verified stays clean; a stamp a human removes or changes resets the count to zero.
Findings: what the engines caught
Every detector writes here. This is the work queue, not a score.
The three states are disjoint and sum to 527. Cleared is NOT confirmation the fix works; only verified means it was deliberately re-probed after a fix.
A repeat is a finding whose class (identity, taxonomy coverage, contract completeness…) already had a prior finding — the measure of whether fixes land on the class or just the instance. Steady state is the 30-day rate at zero: new findings only ever open new classes. This number is expected to start ugly; publishing it is the point.
By failure mode and state
How long the open ones have been open
A tall bar on the right is the treadmill this page exists to end: detection outrunning remediation. An empty right bar just after a stale-findings sweep means old entries were re-probed and cleared, not remediated.
Recently cleared
Consumer findings from Raven
Defects filed against this service by stellar-raven, its largest agent consumer, from that project's own evaluation battery. These carry more signal than our internal detectors because the answer key is not ours.
Their answer key, by status
A record can read reported-upstream on their side while our linked issue is closed: we shipped the fix and their re-verification has not run yet. Neither status is allowed to speak for the other.
Defect flow: detector to surface to outcome
Every finding in the ledger traced through the system: which detector caught it, which surface it lives on, and whether it closed. Ribbon thickness is the count; whole-ledger, not a sample.
Read left to right: a detector produces findings, they land on a surface, and they end Cleared, Open, or Verified. A fat ribbon into Open is a surface carrying real debt; a fat ribbon into Cleared is a detector whose class has been closed.
Where open recall misses die
Each open recall finding replayed live and classified at the FIRST stage that fails - mutually exclusive classes with different owners, not a funnel or a sequence.
0 of 0 open recall findings replayed
Row quality: the evidence behind each record
Every project row scores on five facts we either hold or don't: a provenance basis, a date, a source URL, a type, and a link. A low score names exactly what is missing.
Status provenance, strongest evidence first
Deployment fact (sls-079): 903 unknown · 79 mainnet · 2 testnetiWhich network a product is deployed on, as a separate fact from lifecycle status. Populated ONLY from evidence (verified mainnet contract joins, on-chain readings, human-verified operator artifacts); unknown is the honest default and a work queue, never a score. The gap matrix carries the prominent rows to work first.
Most rows rest on site-liveness - a page answered. That is the weakest honest basis we serve, and moving rows up this ramp is the standing data job.
What is missing, across all rows
Every row as one dot. The marked region is the curation queue.
Curation queue: rows whose evidence is thinnest, prominent firstiSorted by evidence score ascending, then by curated prominence, so the most-seen thin rows surface first. Each line names exactly which of the five facts is missing.
The code index: what it holds, how deeply we know it
2,920 curated repos (claimed by a project or a tracked builder) plus a 10,018-row Electric Capital tail indexed for completeness. The charts read over the curated index only - mixing the tail in made the curated index look unscanned when it is not.
Repo score distribution, curated indexirepoScore (0-100) grades freshness, traction and builder authority. A long low tail is EXPECTED in an open ecosystem - hackathon one-offs and early experiments are real code references worth indexing; the score is what keeps them ranked below production repos.
Commit activity, curated indexiDerived from each repo's last commit: active (<=90d), slowing (<=1y), dormant (older), archived, or unknown (no commit date held - not knowing is its own state, never counted as dormant).
Languages
How deeply we know each layeriThree different jobs with three different denominators. Depth scanning aims at the whole curated index. Knowledge notes are a hand-curated research layer being built over the highest-scored repos - a small number is early progress, not missing homework. Mainnet joins are deliberately strict: only a verified on-chain attribution counts, so the number grows slowly and every unit of it is proof.
Curated repos with a code-depth reading (entry files, symbols, SDK usage). The scan waves aim at all of them.
Hand-curated dated facts with sources, written over the top-scored pool one repo at a time. An enrichment layer under construction, newest additions first.
Deployable contracts with a PROVEN on-chain attribution. Strict by design: absence is absence of a join, never proof of disuse.
Notes pool, honestly split: of 408 pool repos, 151 carry dated facts, 195 were examined and yielded nothing durable (judged, recorded internally), and 62 are still unexamined. A judged repo is not a gap.
Plus the Electric Capital tail: 7,935 of 10,018 rows scanned opportunistically as budget allows - indexed for completeness, no coverage target attached.
Highest-graded repos
Lessons and research
Every recurring defect class was written up when it was found, and every human-verified correction carries a committed receipt. These are the documents behind the numbers above.
Lesson write-upsiEach file records defects found on one day: what broke, the root cause, and the invariant or probe added so the class cannot silently return.
Correction receiptsiA human-verified status change commits its evidence: the URL fetched, the time, response identity headers, and the exact markers on the page that decided the verdict. Re-run the capture to diff what a page says now against what it said then.
Audits
Trends
Daily history appended by the eval pipeline and committed, red days included. Battery probe counts rotate with the daily banks, so the pass line moves by design; the failure line and the ratchets are the signal.
SCF funding cross-check
No project overstates or understates SCF membership, at the project level OR the round level, against the fund's own directory.
- 310 matched records checked against 500 SCF projects
- 0 overstated / 0 understated at project level
- 0 round-level overclaims across 309 verified claims
Contract honesty probe (Engine E)
Documented params do something, undocumented values are rejected. The contract a stranger hits behaves as written.
- 0 param(s) documented but silently ignored
- 0 param(s) accepting values the spec forbids
- spec 1.9.1 at measurement; re-run on every deploy
Curated canonical repos resolve
Every repo we call authoritative is indexed and carries code signals — a curated name that matches no row silently degrades the query it was written for.
- 0 absent (curated name matches no row)
- 0 indexed but no code signals — invisible to code-evidence ranking and to the tier gate
Computed values reach a serving path
Every field our machinery computes is read by something that shapes an agent's answer — a value nothing consumes cannot change what anyone is told.
- a script or a test does not count as consumption — that is how codeProofTier passed for months while only a report called it
Guard lanes that actually run
Every automated lane in the repo has completed a real run recently — a guard that never executes is a promise, not a check.
- FAILING: api-drift.yml — 2 failed run(s) since its last green 2d ago — it is trying and losing, not idle; dies at "Field-population guard (values arrive, not just shape)"
- FAILING: workflow-health.yml — 5 failed run(s) since its last green 1d ago — it is trying and losing, not idle; dies at "Mirror drift (scout-mcp)"
- judged against the workflow file as it stands: runs from a since-edited or never-merged version are ignored, so a fixed lane stops being red
- a lane that exits 1 to report a finding is the guard working, and is not counted here
Type errors in scripts/
The scripts that write to the production database are type-checked, and the backlog of known errors only shrinks.
- identity is file + error code + message, without line numbers, so an unrelated edit above an error does not read as a regression
- a count of 0 here means the ratchet is finished and the guard can become a plain tsc gate
Served values a consumer can date
Every value we serve carries a date that covers IT — or says plainly that it cannot be dated. A value wearing a neighbour's timestamp is a wrong answer with a citation.
- explainRepo.answerAsOf is the pattern: NULL for a DeepWiki answer, because DeepWiki exposes no index date and inventing one would make an unknown look measured
- scoping cuts both ways: verifyClaim's confidence.ageDays dates the numbers inside confidence but NOT the root-level verdict beside it — so verifyClaim.verdict is carried as named debt, not excused; an admission only counts when the description speaks to datability itself ('no index date'), never bare 'null' or 'unknown'
- limit: a scope holding one date is treated as dating every value in it, so this catches the sharper shape only — a value with NO date in scope while other objects in the response carry dates
SCF-funded projects served
Every SCF-funded project (round-badged, so provably funded) is in the directory an agent searches.
- 0 SCF-funded projects the directory does not serve
- DefiLlama: 0 missing of 42 Stellar-listed
- absent entries are human-reviewed SEEDS, never bulk-created
Code-depth calibration
Repo depth grades agree with independent code analysis where both exist — on a sample large enough for the rate to mean something.
- 0 disagreements
- two graders, recorded per row and never pooled silently: DeepWiki for the few answer-key repos it has indexed, Grok (agentic, reads the repo) for the rest
- the binding constraints are EXTERNAL: DeepWiki's index covers ~4 of the 86 long-tail answer-key repos, and the Grok balance exhausted mid-run (HTTP 402) after 9 verdicts — topping it up or DeepWiki indexing more of the key is what grows n
- sample meets the 20-repo floor
Consumer interlock (Raven)
The #1 consumer's discovery index tracks our contract, checked from OUR side too, with grace for their re-baseline cadence.
- 4 op(s) lagging within the 10-day re-baseline grace window (expected)
- 0 op(s) missing beyond grace
- contract 1.9.1 at measurement
Recall floors (Engine A)
Generated known-item probes per bucket stay above their red-line floors, recall can't silently erode.
- all buckets above floor
Real-demand OK-rate (Engine D)
The queries real consumers actually sent are replayed live and keep answering at or above the committed floor. A miss on real demand outranks any synthetic finding.
- 9,254 real-consumer calls in the 14-day window
- 2 queries missing today, the standing fix queue
- floor 80% is a ratchet from a prior reading, it may only move up
Golden retrieval eval
Known-true questions (answer key derived from the canonical directory) keep passing after every ship.
- full pass
- 2 N/A by design (live-source questions the static corpus doesn't answer)
- re-run on every production deploy + weekly
Corpus hygiene (S5-S8)
The research corpus stays clean: junk URLs, broken titles, staleness and mirror drift are swept weekly, and a sweep that returns no reading counts as unknown, never as clean.
- s7 coverage: 2 source(s) structurally undateable (declared)
- all sweeps clean
Consumer path (through Raven)
The canonical questions answer correctly through the REAL Raven gateway: routing, our op, the response envelope and coaching, not just our direct API.
- every gradeable golden question answers correctly through Raven (47 checked)
Improvement ledger
Every quality detector's findings land in one tracked backlog by surface. A backlog is fine; a HIGH-severity finding neglected past 30 days is the failure, that's this row's red line.
- open by surface: consumer 5 · retrieval 0 · code 0 · directory 0 · scf 0 · contract 0 · corpus 0
- 527 tracked · 0 in-wave · 7 verified · 514 auto-cleared (detector stopped flagging)
- closure basis: 7 verified deliberately · 216 cleared on a live re-probe that PASSED · 298 cleared only because a detector stopped reporting
- that last group is the re-probe backlog, not a result: a detector going quiet is indistinguishable from a gap nobody asked about again. Spot-check 2026-08-31 — kutana, etesia and octopos sit in it, are still absent from the directory, and each carries SCF round badges.
- no high-severity finding neglected past 30 days
- all detectors reporting within 10d, every open finding is a confirmed one
Response-shape opacity (build-enforced)
No silent opacity: every response object declares its shape. Grandfathered open maps are counted here and the count may only fall.
- 0 response objects still return an undeclared shape
- the ratchet fails the build if that number rises; it may only fall
- a NEW operation shipping without a declared shape fails the build outright
- scripts/contract/check-schema-opacity.ts, wired into contract:check
Match labelling (build-enforced)
Every q-taking operation declares HOW it matched, so a caller can tell an exact hit from a semantic neighbour. Enforced in CI.
- 4 q-taking operation(s) tracked, 0 unlabelled
- a new q-taking operation shipping without a match label fails the build
- scripts/contract/check-honesty-layer.ts, wired into contract:check