Method
How Subfeel works
A public version of the instrument’s rules. Every score published anywhere on a Subfeel register follows this sequence.
- 01
Nine feelings, nominated by neuroscience and kept by measurement
The instrument reads joy, interest, affection, desire, sadness, anger, fear, disgust and surprise. Seven of the candidates come from the primary emotional systems described in affective neuroscience; disgust and surprise join as sensory reflexes review text is full of. That grounding does less work than the word neuroscience suggests: it supplies an exclusion rule, not a proof. Serious accounts disagree, and if the constructionist camp is right these are concepts shared by English speakers rather than universal circuits, which would limit claims across languages. So membership is decided by measurement instead: a feeling keeps its place only while independent readers can agree about it, and one boundary has already been rewritten after a blind test contradicted the author's own ruling. Subtler concepts such as cosiness, awe or gratitude are composed downstream rather than asked for directly, because they are blends and a rater cannot weigh a blend by eye. The nine are independent flags, not opposites. A good horror game evokes fear and joy at once.
- 02
A feeling counts in two channels, and the reader must know which
A feeling reaches the reader through two channels. The writer says "I laughed the whole way through" and never says happy; the joy is implied, and the instrument reads it. So a feeling counts when it is genuinely there, whether the words state it or it sits between the lines, and every human label records which channel it came from. What never counts: a verdict. "Great game" with no feeling behind it is a No, because enthusiasm gives off a warm glow and untrained readers turn that glow into ticks on every positive feeling at once. The one unforgivable error is claiming the words state a feeling they do not. Reading the implied channel honestly, without inventing what is not there, is the whole instrument.
- 03
The humans behind the ground truth must qualify first
Model scores are only as good as the human labels behind them, so labelers qualify before their labels count. Each rater studies a written guide, then passes a scored exam whose answer key cites a written rule for every answer. Qualified raters then have to agree with each other at a pre-registered statistical bar before any of their labels become training or test data. When agreement fails, the guide gets revised, not the bar.
- 04
Traps catch what is not a person
Labeling rounds carry embedded integrity checks: texts a reading human cannot answer wrong, and one text containing an instruction only an AI assistant would obey. Across our paid crowd studies, roughly one in five participants failed these checks and was excluded. Timing floors, per-participant device tokens and an append-only submission log close the remaining forgery routes. What survives the filters is human judgment, demonstrably.
- 05
The model is small, local and domain-tuned
The reader is an open-weight model fine-tuned on qualified labels, running on local hardware. On its validated text types it scores far above the standard open emotion classifier; off them, it goes blind in exactly the same way that classifier does outside Reddit — we measured our own blindness at 1% joy detection on health-app raves and published it. The lesson includes us: emotion classifiers learn corpora, not feelings. That is why every number ships with the text type it was verified on. The current engineering program trains for cross-register generalization under a written law: a detector may claim to read a feeling only after blind human judgment confirms it at 0.8 precision on text types it never trained on, with the statistics to back it (fifty items per type, confidence bound included). Tested where it trained counts for nothing.
- 07
The definitions are the instrument
The sharpest result we have: a frontier model asked to find affection in health-app reviews scored 43%. Given one paragraph of our definition, written from a founder's blind rulings, the same model scored 100% on the same texts (26 of 26, human-judged). The weights did not change; the definition did. That is what this project actually builds: operational definitions of nine feelings, ruled by a consistent human judge, tested where they were never fitted, and portable to any model. Models come and go. The codebook is the instrument.
- 06
An independent adversary rejects our own proposals
Every scoring rule, handbook and key passes an adversarial reviewer that is structurally separate from the builder, computes its own numbers, and has rejected our own proposals repeatedly before approving a version. Design decisions are pre-registered before data arrives, founder intuitions are tested blind and overruled when they fail, and the failures are published alongside the wins. The check only ever finds mistakes; it never fixes them itself. That separation keeps it honest.
Scores are never for sale
A score derived from public review text is never suppressible, hideable, or removable for money, no matter who asks. Privacy is only ever sold on data we could not otherwise reach in the first place. See for brands for the sealed, private-corpus option. The open register and the private scoring service stay permanently separate.