subfeel

FeelBench

The implied-feelings benchmark

Most feeling in real reviews is implied, not stated: we counted two explicit first-person anger statements in 18,462 texts. FeelBench measures whether a model can read the implied channel: 48 audit-certified review texts across four domains, nine feelings, scored against human-ruled gold with a single plain prompt, zero-shot, macro-F1 over the gold-active classes.

modelmacro-F1ticks (gold: 107)
Fine-tuned 14B (Subfeel), one desk Mac0.545106
Gemini 3.1 Pro preview (Google)0.498105
GPT 5.6 Luna Pro (OpenAI)0.49695
Kimi K3 (Moonshot)0.48099
Claude Opus 5 (Anthropic)0.42065
Claude Fable 5 (Anthropic)0.41063

Same 48 items, same prompt, same scorer for every row. The scorer reproduces the committed historical answers exactly before any new model runs. New runs made August 2026 via OpenRouter; answer files retained for independent rescoring.

The caveats, stated first

The gold is crowd-majority with a founder-audited inflation of 1.0 to 2.1x, and the 14B was fine-tuned on that label distribution, so part of its margin is calibration rather than comprehension. CORRECTION AND FINDING (2026-08-18): this benchmark is our home ground, and the honest reading changed. Within its training registers the fine-tune wins; measured OFF those registers it collapses below the frontier models (for example, 1% joy detection on five-star health-app reviews). The general lesson, which includes our own model: fine-tuned feeling classifiers learn corpora, not feelings. We no longer claim to beat frontier models at reading feelings; we claim measured precision on named text types, and we publish where we are blind. That is also the finding about the frontier generation closed the tick-calibration gap and still landed five points short, because the remaining distance lives in ruled, labeled boundaries that scale does not buy. n is 48; pairwise gaps between neighbouring rows are inside noise; the gap to the previous frontier generation is not.

Think your model does better?

The harness is a single script: same prompt, temperature 0, JSON answers, committed scorer. Write to per@subfeel.com and we will run your model, or send you the items and scoring so you can refute us properly. Refutations get published on this page.