FeelBench
The implied-feelings benchmark
Most feeling in real reviews is implied, not stated: we counted two explicit first-person anger statements in 18,462 texts. FeelBench measures whether a model can read the implied channel: 48 audit-certified review texts across four domains, nine feelings, scored against human-ruled gold with a single plain prompt, zero-shot, macro-F1 over the gold-active classes.
| model | macro-F1 | ticks (gold: 107) |
|---|---|---|
| Fine-tuned 14B (Subfeel), one desk Mac | 0.545 | 106 |
| Gemini 3.1 Pro preview (Google) | 0.498 | 105 |
| GPT 5.6 Luna Pro (OpenAI) | 0.496 | 95 |
| Kimi K3 (Moonshot) | 0.480 | 99 |
| Claude Opus 5 (Anthropic) | 0.420 | 65 |
| Claude Fable 5 (Anthropic) | 0.410 | 63 |
Same 48 items, same prompt, same scorer for every row. The scorer reproduces the committed historical answers exactly before any new model runs. New runs made August 2026 via OpenRouter; answer files retained for independent rescoring.
The caveats, stated first
The gold is crowd-majority with a founder-audited inflation of 1.0 to 2.1x, and the 14B was fine-tuned on that label distribution, so part of its margin is calibration rather than comprehension. CORRECTION AND FINDING (2026-08-18): this benchmark is our home ground, and the honest reading changed. Within its training registers the fine-tune wins; measured OFF those registers it collapses below the frontier models (for example, 1% joy detection on five-star health-app reviews). The general lesson, which includes our own model: fine-tuned feeling classifiers learn corpora, not feelings. We no longer claim to beat frontier models at reading feelings; we claim measured precision on named text types, and we publish where we are blind. That is also the finding about the frontier generation closed the tick-calibration gap and still landed five points short, because the remaining distance lives in ruled, labeled boundaries that scale does not buy. n is 48; pairwise gaps between neighbouring rows are inside noise; the gap to the previous frontier generation is not.
Think your model does better?
The harness is a single script: same prompt, temperature 0, JSON answers, committed scorer. Write to per@subfeel.com and we will run your model, or send you the items and scoring so you can refute us properly. Refutations get published on this page.