What a reading is allowed to claim, and the standard it has to meet before it is published. The pipeline that produces it is a separate page.
A reading is a claim about a public, and every claim here carries the apparatus that would let someone else disagree with it precisely.
The front door puts the difference in two sentences: social listening, sold on, is closer to surveillance, and an insight carried back into the room it came from is the root of acquaintance. Here is what stands behind them. A population is read, the reading goes to whoever can pay for it, and the people counted never see the count — and nothing in the method requires that. The method does not change when the reading is published back instead of sold on; who ends up holding it does. So the corpus, the instruments and the daily sheet are open, at k ≥ 20, to anyone.
Four rules govern anything published on this host. They are not style preferences. Each one has a gate behind it, and one of them runs on every deploy.
The query, the window, the run, and the code path that produced it. A number without a provenance line is not publishable. This is the rule that makes the rest possible, because it is what turns a disagreement into a reproducible one.
A number a human read off a surface is heard. A number the pipeline
computed is counted. Most current dial numbers are heard, and
the surfaces say so rather than rounding the distinction away. The rung a number files
at follows from the mark: what a person read off a surface files as
reported, what the pipeline computed files as thin or
measured by the depth of its baseline.
Gaps, failed harvests and thin coverage appear at the same size and in the same voice as the wins. A source that returned nothing says so, in place. A page that only shows what worked is a page that has been edited into an argument.
Surfaces report counts, deltas, timestamps, dial values and resolutions. No analyst voice, no adjective doing the work of evidence. Specimen copy simulates the harvest, never the analyst. A lexical gate runs on every deploy and refuses the known offenders.
The claim a reading makes is not "trust this number". It is "run this and get the same number". Every published reading prints the query beside the figure so the second thing is possible.
The final step of reproducing a reading is a diff against what was published. Reproduction that stops before the comparison is a tutorial, not a proof.
This is also the contribution route that is open today. A disagreement about arithmetic is a factual dispute rather than an authored work, so it needs no licence from you and no settled terms from us. If your diff disagrees with ours, send it.
A public is described here by what it does and the codes it runs on: its lexicon, what it treats as trying too hard, which references mark membership, what it turns away. Age bands and income brackets describe the same people to somebody else, and a reading can carry both; the rule set is the part it is built on.
This follows from the gate in the mission rather than from taste. Demographics describe a group to somebody else, which is what makes them the natural unit for a market report. Rule sets describe a world to the people inside it, which is what makes a reading survivable when it is handed to its own subject.
The practical consequence is that a persona here is a rule set plus its evidence: every rule in it points at the posts it was read from, and a rule that cannot is filed as a hypothesis until it can.
The standard above is abstract. These are the places it is carried out, and they are the pages to read if you want to argue with the method rather than about it.
fig 1 . where the method is worked out
| Record | What it carries | Route |
|---|---|---|
| Audience Decode | The instrument record: the premise, the six layers, the formulas, the measurement ladder, and what it does not claim. | /decode/ |
| The field | The working instrument over a worked example, with the console and the report. | /decode/field/ |
| The directory | Every kind of data the house knows about, with usability as a status tag rather than a filter. | /data/ |
| The current reading | The climate desk: the nine streams and what they currently say. | /desk/ |
The measurement ladder in the Decode record says how far the instrument has been validated, rung by rung. Read that before the formulas.
heard rather than
counted. That is stated on the surfaces that carry them rather than fixed by
relabelling.The section above says what the method does not claim. This is a run where that was the whole result. On 26 August 2026 the instrument was pointed at one segment, skincare actives, over one home room and 23 control rooms. One of the six Decode layers filled. The other five abstained, and each one says what it would take to fill it.
fig 2 . six layers, one segment, 21 days
63 terms are dense in this room and near-absent from the 23 control rooms, every one over 20 distinct accounts. The set is a product shelf and a symptom vocabulary, not a coinage vocabulary: no in-group slang, no status markers, no algospeak. This room names things in the world rather than naming itself. TWO CAVEATS THAT BITE. The room has no comments harvested, so this is a reading of post titles and selftext only, and titles are where a room is most formulaic. And moderator-defined post flair arrives inside the title in brackets, so terms like 'routine help' and 'product question' are the room's filing system rather than its speech; they are now classified as venue-flair by the collision stamper and excluded from any sociolect claim.
What it would take. Comments, and a reference extractor. Two blockers, and the first is the one that matters. THIS ROOM HAS NO COMMENTS HARVESTED, so the co-presence measure is DEGENERATE rather than null: every thread holds exactly one row across all 585 threads, so every row is a top-level post and no two accounts have ever been observed in the same conversation. Co-reference, reply structure and every other canon signal are unavailable by construction. Nine of the 24 rooms are in this state. Separately, items.url is the item's own permalink, so a shared-URL test over url_fp cannot work either: every value is distinct by definition, and reporting that as 'no shared references' would be an artifact dressed as a finding. The first draft of this script did exactly that.
What it would take. Stance detection over the corpus. The house has run stance exactly once, on 2026-08-22 over the interpreter corpus, as a one-off. There is no routine, no radar.stances table, and no scored run. Until there is, any sentence here about what this room values would be written by a reader rather than measured, which is the copy law's exact prohibition.
What it would take. A second platform. This segment is read on 1 venue, all of it Reddit. Cross-platform shape is a comparison and one venue is not a comparison. The fediverse lane opens and #skincare is harvestable there, so this is the layer closest to unlocking.
What it would take. An in-group and out-group marking method, which the house has not built. Reading who counts as inside means reading correction, gatekeeping and repair, all of which live in COMMENTS. The comment corpus is gated on T2 and the retention window on T1, both unruled.
What it would take. Time. The corpus covers 21 days across 4 weekly windows. Drift, half-life and obtrusiveness all need a slope, and a slope over four points is a drawing. Arctic Shift backfill (T5) is the route to the history this needs.
generated 2026-08-26 · from the segment-read run · not hand-edited. Five of six layers say nothing, and each says why. This is a complete run of an incomplete instrument, not a partial run of a complete one. Treat the one filled layer as the whole finding.
The file prints what the run does not claim, and the first item is the one that matters: it does not claim the segment has no values, no boundaries and no trajectory. It has all three. The instrument cannot see them yet, which is a different sentence and the only one the evidence supports.
The corpus behind this run is described on the mechanism page, which is where the two blockers are counted across all 105 rooms rather than in this one.
The sociolect question is whether a room has vocabulary that is dense in it and near-absent everywhere else. Run on 26 August 2026 over eight weekly windows, 24 rooms and 8,748 distinct terms, the answer was that no room did. The run is kept here because a measurement that returns nothing and says why is evidence about the instrument, and discarding it would leave only the runs that found something.
fig 3 . 8,748 terms, and what the 83 survivors turned out to be
| what the 83 turned out to be | terms |
|---|---|
| general-english | 59 |
| not-in-general-english | 9 |
| venue-ritual | 6 |
| venue-named | 5 |
| venue-flair | 4 |
| term | home room | accounts | specificity | windows |
|---|---|---|---|---|
| sunscreen | reddit-skincareaddiction | 55 | 0.92 | 4 |
| moisturizer | reddit-skincareaddiction | 51 | 1.00 | 4 |
| spf | reddit-skincareaddiction | 36 | 1.00 | 4 |
| cerave | reddit-skincareaddiction | 32 | 1.00 | 4 |
| 2000s | reddit-decadeology | 30 | 0.83 | 7 |
| coastline | reddit-explainlikeimfive | 25 | 1.00 | 1 |
| niacinamide | reddit-skincareaddiction | 23 | 1.00 | 4 |
| tretinoin | reddit-skincareaddiction | 22 | 1.00 | 4 |
| vietnam | reddit-askoldpeople | 22 | 0.81 | 1 |
generated 2026-08-26 · from the sociolect run · not hand-edited. Dictionary: the system word list, 234,456 entries.
The last nine are the terms that cleared every filter. The class name says only that they are absent from one dictionary file, and the table says nothing beyond what was counted about each.
None of the nine is a coinage. They are common nouns, two ingredients, a brand, a decade and a country: words missing from one 234,456-entry spell-check list rather than words a community made. The run's own headline says the same thing, that no distinctive shared vocabulary was detected in any room.
Read that as the instrument's answer over this corpus rather than as a statement about language. The 24 rooms are general-interest Reddit rooms in advice, confession and explainer formats, and a question format is a different object from a subculture. Forty-two of the 83 candidates come from one room.
The Berkeley Protocol on Digital Open Source Investigations asks anyone collecting at scale to state three things in writing: the purpose the collection serves, why it is necessary for that purpose, and what is not collected. It is explicit that the principle favours itemized, manual collection over bulk. This house is bulk collection across eighteen ingesters, so the burden is ours. This section is it.
Purpose. To make social rule sets visible, nameable and changeable by the people living inside them, which is the mission. Every published output of this house is a statement about a set. There is no page, product or deliverable whose subject is a person.
Necessity, which is the part usually left unsaid. A social rule is not observable in an individual. One person using a term tells you that one person used it, not that the term is a norm. A norm is a property of a population, and the smallest object that can carry one is a group. So the floor of twenty distinct accounts behind every published cell is doing two jobs, and this house has only ever claimed the first:
fig 4 . what the floor is for
| The floor read as | What it does | State |
|---|---|---|
| A privacy floor | Stops a published cell being narrowed to a person | stated |
| The unit of analysis | Is the smallest cell that can carry the claim being made | stated here |
A cell below the floor is not a risk the house declines to take. It is a claim the house cannot make, which is why a reading abstains rather than publishing a small number with a caveat beside it. The caveat would not fix it: the number was never evidence.
Three things follow, and each is why the line is drawn at what leaves rather than at what is read. You cannot count distinct accounts without counting accounts. You cannot detect a community without an account-level edge list, because a community is defined by who overlaps with whom. And you cannot know a cell is large enough to publish without holding the data the count is performed on. The alternative the Protocol nominally prefers, a hand-picked sample, produces a different object: one whose selection criteria are unstated and whose bias is therefore invisible.
Proportionality. The line is drawn at egress: read individuals, publish only crowds. What follows is enforced rather than intended.
Two places where this is not yet satisfied, printed here rather than left for a reader to find.
The people this applies to have never interacted with this house. They posted in public rooms and do not know it exists, which is the reason the burden binds at all. The answer is that nothing is ever published about one of them, only about populations they are part of. Someone who does hand this house data has a stronger claim to a stated clock, not a weaker one, and that clock is ruled before the first seat is issued rather than after.
the method is held to the standard above on this host only. commissioned work is out of scope here by house rule and is not described on any gossip surface.