Ptakha >> Sutskever

Ptakha >> Sutskever

Ilya Sutskever gave an interview. I think it's an important one. Ilya is a co-founder of OpenAI, one of the three people who literally built modern AI. Two years of silence. In November he talked. Here's what matters:

1/ The scaling era is over. The research era has started.

The trend: 2020–2025 was the age of scaling. More data + more compute = smarter model. Scaling laws worked. Companies weren't doing research, they were scaling up things that had already been invented. But the internet is finite (and that's what the models trained on). In 2026 and beyond, breakthroughs need new ideas. And there are fewer ideas than there are companies building foundation models.

Why it matters: - The window is open. While Google, OpenAI and Anthropic hunt for the next breakthrough, this is the best time in 3 years to build products on the models we already have. - Whoever first ships a product where the LLM makes the UX several times better wins.

2/ Benchmarks ≠ economic value

The trend: models post great numbers on evals and then fail basic real-world tasks. I run into this constantly. First the agent does something the average human couldn't, then (literally in the next message) it loops and says the same thing three times. Why? Companies optimized for benchmarks. What we got are excellent test-takers and homework-solvers that work badly on real tasks.

Why it matters: - Don't trust benchmarks when picking a model for actual work. - Build your own evals for every task, and compare models on those. - Easier said than done. We (and every team actually shipping something with LLMs) ate a lot of shit on evals this past year. It won't be easy.

3/ Pre-training commoditizes, post-training differentiates

The trend: all foundation models are trained on roughly the same data. Base capabilities don't differ much. The difference shows up in post-training: RLHF, domain-specific fine-tuning, and so on.

Why it matters: - Don't try to build "the best foundation model". Take open source: Qwen, DeepSeek, Gemma are already at SOTA level for many domains. Put your money into post-training on your own data. - But! Good post-training needs infrastructure, for example a decent online environment for RL. That's where the investment goes.

4/ Vertical > horizontal

The trend: there will be a lot of narrow AI companies. Legal, medicine, support, finance. It's already happening: Harvey (legal AI), PathAI (medical diagnostics).

Why it matters: - The first one to assemble the puzzle in a specific domain (unique data + domain-specific evals + post-training infrastructure) can monopolize that vertical. - Literally: beat ChatGPT at support. Or at legal documents. General-purpose models will stay, but business value capture shifts into the verticals.

5/ The generalization gap

The most important line of the whole interview:

Models somehow just generalize dramatically worse than people.

A kid hears a new word 2–3 times and starts using it correctly. Models need thousands of examples. This isn't a scale problem. It's a fundamentally different mechanism, and we don't understand how it works.

Why it matters: scaling is done, and nobody knows how long the wait for the next breakthrough is. Ilya's estimate: 5 to 20 years. A 4x spread speaks for itself.


P.S. I deliberately left out everything about AGI and AI safety. Both topics are super-hype. We won't hit them any time soon, nobody knows how any of it plays out. More speculation than practically useful insight.

P.P.S. Two interviews came out a week apart. Ptakha on Dud's show: 6M views. Sutskever on Dwarkesh: 1M. Interesting world we live in.

Stay tuned 🥷🥷🥷


More takes — @tldrdaniel
Written by Daniel Levinishnikov