AI Change Tracker

Fiction.liveBench

https://fiction.live/stories/Fiction-liveBench

Kind
benchmark
Maintainer
Fiction.live
Feeds axes
long_context
Measures
Whether a model actually understands a very long document end to end — who knew what, when things happened — rather than just claiming a big capacity on the label.
Does not measure
Long business documents or code specifically (it uses fiction), finding a single planted fact, speed on long inputs.
Known issues
Built on fan fiction, which is a narrow style of text; question quality varies; relatively few questions at each length.
Trust
medium
Last reviewed
2026-08-01

Every model advertises how much text it can take in at once — hundreds of pages, in some cases. This test checks whether the model actually understands text that long, by feeding it long serialized stories and asking comprehension questions that require holding the whole narrative in mind: who knows what, what happened before what. That's much harder than finding one planted sentence, which is what vendors' own demos usually show.

How to read a score: this site records the score at the longest length bucket (roughly a 500-page book). The shape of decline matters more than any single number — a model that scores 95% on short text and 60% at full length has a real working capacity far below what the label says. When two long-document tests disagree, trust neither and read both.

← All methods