About
A personal, continuously updated site that answers four questions with sources attached: what changed in the AI industry this week, which model is currently best for each of a set of named use cases, how those judgments are made — and, above all of it, what the changes mean and enable: the strategic read and the opportunities it opens.
The site is a set of views over one structured dataset in a git repository. A scheduled pipeline fetches lab blogs, leaderboards, and curators; classifies developments into events; appends scores; and drafts the weekly strategic synthesis. Every number and event carries a source URL and an as-of date. Astro renders it statically; Vercel deploys on every push.
How to read the numbers
Scores are numbers from named sources — a score marked L is lab-reported and not yet independently reproduced; treat it as a claim. A marks aggregator-measured values. Each source's strengths and blind spots are documented on Methods — score popovers link there.
Three voices are kept strictly separate and styled differently. Data (scores, events) sits on plain white. The owner's judgment — verdicts and theses — appears in white cards with a strong navy edge, written and changed only by the owner. Machine-drafted analysis — event implications, the weekly Signal, opportunity proposals — appears in navy blocks, always labeled, and is a proposal for the owner's judgment, never a substitute for it. The numbered badges on stories are significance, 1–5: how much a development changes what matters (darker = bigger), assigned against a written rubric.
Map of the site
- Home — The two-minute daily read, in a fixed order: what needs your decision, top headlines, what’s rising, current verdicts, thesis drift. Stop after the headlines on most days.
- Signal — The strategic layer, written weekly: what the week’s events mean read together, which positions gained or lost ground, and proposed opportunities. Meaning, not news.
- Feed — Every event, newest first, filterable by topic, lab, significance, and date. The raw stream Home and Signal are distilled from.
- Compare — The data: a sortable matrix of entities × evaluation axes, every cell tap-able for its sources and dates.
- Verdicts — The owner’s current pick per use case, with rationale, evidence, and history. Judgment, clearly separated from data.
- Theses — The owner’s falsifiable positions, with confidence history and evidence drift as events support or challenge them.
- Bets — Other people’s public predictions — forecasters, researchers, markets — with resolution criteria and calibration track records.
- Methods — Why each benchmark and leaderboard is trusted: what it measures, what it misses, and how to read a score. The epistemology page.
- Changes — The audit trail: every significant event, score update, verdict change, and confidence move, interleaved by date.
- Digest — The weekly email-style recap mirroring Home’s reading order — for catching up after a week away.
Built from a design document (docs/design.md in the repo) by Claude Code.
Events RSS: /feed.xml.