How a row gets from a public surface to a number on the desk, which stages are automated, and which are a person doing it by hand.
Five stages. One is fully machine, two are a person, one is both, and one is a gap, printed here at the same size as the four that run.
fig 1 . the pipeline, stage by stage, with what runs today
| Stage | What happens | State |
|---|---|---|
| 1 . registry | A source is named, given a kind, and given a status. Unusable sources stay on the list with the reason. | machine + hand |
| 2 . harvest | Rows are pulled on a schedule. Failures are recorded as rows in their own right, so a dark source is visible rather than absent. | machine |
| 3 . banding | Rows are grouped into the categories a reading is made of. This is the authored step and it is a person's judgement, not a classifier. | hand |
| 4 . streams | Banded rows become the nine streams the desk reads from. | not built |
| 5 . the dial | A stream becomes a number with a spec line and a provenance mark. | hand, marked heard |
Stage 4 does not exist, so it is drawn off the path rather than on it, and stage 5 runs by hand as a result — which is why the marks in section 03 matter more than they otherwise would. The figure is authored: unlike the run evidence further down this page, no run verifies it.
Organisations, public accounts, feeds, archives and public APIs. Never an individual, and never a closed space. The full census is the directory, which lists every kind of data the house knows about with usability as a status tag rather than as a filter.
The directory lists the sources that did not work as well as the ones that did. That is deliberate and it is the more useful half: a list of only what worked is a list that makes every future reader re-derive the same dead ends.
A source suggestion is one of the contributions that is open today. It needs no licence from you, because a URL and a factual status tag are not an authored work.
Most numbers on the desk right now are heard, not counted. The surfaces say so, because a mark that is only applied when it flatters is not a mark.
counted means the pipeline computed it and the query is printed beside it.
heard means a person read it off a surface and wrote it down. Both are
evidence. They are not the same evidence, and merging them would delete the only signal a
reader has about which is which.
Confidence is the rung a number files at, and the rung follows the mark rather than
crossing it. A counted number files as thin under a 14-day
baseline and measured at or over it. A heard number files as
reported: dated, and not checked by us. No heard number reaches
measured, which is what the distinction costs and the reason it is worth
keeping.
Of the nine streams, three can be counted cheaply and four are heard by their nature. A stream that is heard by nature is read rather than tallied, and it carries the mark that says so; the mark is what lets a reader weigh it.
not built . stated rather than implied
The nine streams are computed nowhere in the pipeline today. The
aggregation step emits the domain, the domain system, the source, the corpus and the
migration, and no stream at all. Every stream figure currently on the desk was produced by
hand and carries the heard mark.
There is no feed on this host, and the real-time view is not built. The domain that will carry it serves the daily cover and its dated archive, which is a page a day and not a live view.
The contribution surface is partial. Four kinds of contribution are open; the one the corpus is made of is held until the contributor terms are settled. The reasoning is on the contribute page.
This section is the one most likely to be out of date, and it is the one to check first if something here reads as inconsistent with what a surface is showing you.
Two gates stand between the working tree and this host, and both refuse rather than warn.
A lexical check for analyst voice on surfaces that are supposed to report the harvest. It cannot see a dramatised number, only a known offending word, and the page it prints says so rather than claiming more.
The assembled tree is searched for client and cross-brand strings before the host sees a byte. The only thing keeping commissioned material off this host is that it is not in the repository, so the assertion is checked on every deploy rather than trusted.
The boundary gate had a real defect, found on 17 August 2026: it ran before most of the routes were assembled, so it had only ever inspected part of the tree. Two files carrying cross-brand strings would have shipped. It now runs last, after every route is staged. A guard that runs before the thing it guards is decoration, and it is recorded here because the class of error is more useful than the incident.
Section 02 says where the rows come from. This is what they add up to, counted room by room on 26 August 2026 and left here as the run wrote it.
The count exists because of a specific failure three days before it. Three artefacts had been read as findings: plurals taken for coinages, moderator flair taken for vocabulary, a schema artefact taken for absence. Each one was a property of the instrument wearing the clothes of a result. A standing census of what every room can and cannot support is the form of the question that catches those before anyone quotes them.
fig 2 . the corpus, by platform kind
| platform kind | rooms | items | posts only | under k floor |
|---|---|---|---|---|
| 24 | 20,952 | 9 | 0 | |
| steam | 8 | 18,215 | 8 | 0 |
| rss | 31 | 4,465 | 31 | 25 |
| wiktionary | 3 | 4,040 | 3 | 3 |
| youtube | 8 | 2,088 | 8 | 0 |
| hn | 4 | 1,650 | 4 | 1 |
| apple | 4 | 1,600 | 4 | 0 |
| market | 2 | 959 | 2 | 2 |
| youtube-trending | 2 | 799 | 2 | 0 |
| stackexchange | 6 | 712 | 6 | 0 |
| trends | 2 | 477 | 2 | 2 |
| arena | 5 | 428 | 5 | 2 |
| bluesky | 4 | 183 | 4 | 2 |
| newsletter | 2 | 5 | 2 | 2 |
generated 2026-08-26 · from the corpus-health run · not hand-edited. Thresholds: k ≥ 20, 8 windows for a slope.
A room under the k floor has too few distinct accounts to publish any cell at all. A room with no comments cannot carry the two layers that live in comments. Both are properties of what was harvested, not of the communities.
Two numbers govern most of what the instrument can currently say. Ninety rooms hold posts and no comments, and the layers that read canon and boundaries live in comments. One hundred and four rooms have fewer than the eight weekly windows a slope needs. Together they are the reason a segment read fills one layer of six, which is worked through on the method page.
The most careful entry in the file is a null. Where only top-level posts were harvested,
every thread holds exactly one account by construction, so a mean of 1.00 accounts per thread
is arithmetic rather than behaviour. The run stores that mean as null and keeps
the raw 1.00 in a separate field, so a later reader cannot quote it as a finding about how
those communities behave.
this page describes the machine as it runs on 31 aug 2026, including the stages that do not run. if it disagrees with a surface, the surface is newer.