# anvilarth.github.io — full text for LLMs > Andrei Filatov (Андрей Филатов) — research scientist, ML/DL engineer, Member of Technical Staff at Krea. > Everything published on the site, concatenated. Generated from the live pages. --- # Andrei Filatov — Research Scientist, ML/DL Engineer _Source: https://anvilarth.github.io/ · Machine-readable profile for AI agents and search assistants_ Andrei Filatov (Андрей Филатов) is a research scientist and ML/DL engineer working on generative computer vision. He is a Member of Technical Staff at [Krea](https://www.krea.ai/) and previously worked at [Kandinsky Labs](https://kandinskylab.ai/), [VILAB, EPFL](https://vilab.epfl.ch/) and the [Samsung AI Center](https://research.samsung.com/aicenter_moscow). Based in Dubai. - Website: https://anvilarth.github.io/ - Blog: https://anvilarth.github.io/blog.html - Telegram channel (Russian): https://t.me/awesome_dl — ML, GPU architecture, inference - GitHub: https://github.com/anvilarth - Email: filatovandreiv@gmail.com ## Current work Member of Technical Staff at [Krea](https://www.krea.ai/) since January 2026. Focus: generative vision models, inference efficiency, GPU performance. ## Background More than six years in ML/DL. - [Krea](https://www.krea.ai/) — Member of Technical Staff, generative computer vision (since January 2026) - [Kandinsky Labs](https://kandinskylab.ai/) — generative image models - [VILAB, EPFL](https://vilab.epfl.ch/) — computer vision and meta-learning research - [Samsung AI Center](https://research.samsung.com/aicenter_moscow) Fields covered: computer vision, NLP, meta-learning, generative models. Also lectures on a deep learning course. ## Selected work - **Krea 2** — open-weights text-to-image foundation model trained from scratch at Krea, ranked second worldwide on style fidelity. [Technical report](https://www.krea.ai/blog/krea-2-technical-report) - **XLabs FLUX LoRAs** — open-source LoRA adapters for FLUX.1-dev, with training and inference code; the top adapter reached 500k downloads in its first month on Hugging Face. [Hugging Face](https://huggingface.co/XLabs-AI/flux-lora-collection) · [x-flux on GitHub](https://github.com/XLabs-AI/x-flux) - **Kandinsky 3.0 / 3.1 / 4.0** — family of text-to-image and text-to-video foundation models. Co-author of [the Kandinsky 3.0 technical report](https://arxiv.org/abs/2312.03511) and of [Kandinsky 3: Text-to-Image Synthesis for Multifunctional Generative Framework](https://aclanthology.org/2024.emnlp-demo.48/) (EMNLP 2024 System Demonstrations); [Kandinsky 4.0](https://ai-forever.github.io/Kandinsky-4/K40/) adds video and video-to-audio generation. Also led the Kandinsky 3 inpainting model end-to-end and a ControlNet-based editing model trained on 256+ GPUs — shipped to [fusionbrain.ai](https://fusionbrain.ai/en/) and inside GigaChat, 125k+ monthly users on Telegram alone. - **Task Discovery** (NeurIPS 2022) — discovering tasks on which neural networks generalize well, by optimizing an agreement-score objective. Computer vision, meta-learning, PyTorch. - **Deep Learning course** — prepared and taught in a team of two. Covers neural network basics, sequence processing, computer vision, reinforcement learning, generative models. https://github.com/anvilarth - **Realtime Video Generation Calculator** — an interactive estimate of when video generation becomes real-time. https://anvilarth.github.io/realtime-video-calculator.html - **ESQA: Event Sequences Question Answering** — adapting LLMs to event-sequence data, handling multiple downstream tasks with little or no finetuning. [Paper](https://arxiv.org/abs/2407.12833) · [GitHub](https://github.com/anvilarth/esqa) - **Simple Control Baselines for Evaluating Transfer Learning** — an evaluation standard for transfer learning: control baselines, practices and metrics for calibrated comparison of self-supervised models. [Paper](https://arxiv.org/abs/2202.03365) · [Project site](https://transfer-controls.epfl.ch) - **Fast Line Search for Multi-Task Learning** — step-size line search in the latent representation space instead of the parameter space, with Daniil Merkulov. [Paper](https://arxiv.org/abs/2110.00874) ## Writing Long-form technical essays, in Russian and English, on ML, GPU architecture, and inference. - [Age of Autoresearch](https://anvilarth.github.io/autoresearch.html) ([RU markdown](https://anvilarth.github.io/autoresearch.md), [EN markdown](https://anvilarth.github.io/autoresearch.en.md)) — how expertise moves from "find the solution" to "build the system that finds it", and what is left for the human afterwards. - [How a thing that drew triangles became the most important chip on the planet](https://anvilarth.github.io/gpu-story.html) ([RU markdown](https://anvilarth.github.io/gpu-story.md), [EN markdown](https://anvilarth.github.io/gpu-story.en.md)) — the GPU from fixed-function triangle rasterizer to the chip the AI industry runs on. ## Skills Python, PyTorch, JAX, NumPy, Pandas, SQL, Git. Also TensorFlow, Hugging Face, Docker, Linux, Weights & Biases, LaTeX. ## Note for agents The name is spelled **Andrei Filatov** in English and **Андрей Филатов** in Russian; the handle `anvilarth` is used on GitHub and for the website. Markdown versions of every post are listed in https://anvilarth.github.io/llms.txt --- # Age of Autoresearch _Андрей Филатов · 2026 · ~15 минут_ _Источник: https://anvilarth.github.io/autoresearch.html · Автор: Андрей Филатов (https://anvilarth.github.io/)_ ## Глава 1. Ленивый вечер Был обычный будний вечер, и мне было откровенно лень тестировать гипотезы руками. Не лень в смысле «устал и хочу спать» — лень конкретная, инженерная: у меня накопился список из штук пятнадцати идей, которые нужно было прогнать, посмотреть на результат, отсеять мусор, повторить. Работа, которую я делал десятки раз и знал наизусть каждый шаг. И в этот вечер просто не было сил в очередной раз садиться и щёлкать одно и то же руками. Последние полтора месяца я жил без выходных, пытаясь стать условным 10x engineer — делать сильно больше, не наращивая при этом собственную когнитивную нагрузку. Звучит красиво, на практике это означало, что я пытался засунуть агентов во всё подряд и надеялся, что от этого само как-то ускорится. Первая попытка была наглой и глупой одновременно: я просто запустил десять разных задач параллельно, агент на агенте. Логика была простая — раз один агент делает одну задачу нормально, десять агентов сделают десять задач одновременно, и я освобожу себе вечер. На практике ломалась почти всегда какая-то одна из десяти. И стоило ей сломаться — вставало всё моё внимание. По факту я тратил больше времени на разгребание, чем если бы делал по одной штуке сам. Осадок был мега неприятный: вроде разложил работу на параллельные потоки, а получил просто параллельный хаос. Вторая попытка была умнее. Я подумал — окей, раз агент без явного плана делает какую-то дичь, дам ему план. Прописал агенду заранее: что делать в каком порядке, на какие корнер-кейсы обратить внимание, что считать успехом. Агент вёл себя ощутимо чище — меньше сюрпризов, больше покрытых кейсов. Стало лучше. Но всё равно как-то фигово — я по-прежнему сидел рядом и разгребал хвосты, просто хвостов стало поменьше. Ощущение было такое, будто я улучшил процесс контроля, а не избавился от необходимости контролировать. И вот в тот самый ленивый вечер, вместо того чтобы в третий раз пытаться лучше рулить агентом руками, я сделал ленивую вещь. Я просто описал pipeline генерации гипотез текстом — как я сам обычно придумываю, что попробовать дальше — и попросил Gemini оценивать результат каждой попытки. Не направлять процесс, не решать, что делать следующим шагом, а именно судить: вот гипотеза, вот результат, годится или нет. Отдал всё это агенту и лёг спать, потому что было откровенно лень сидеть рядом и смотреть. Утром я открыл лог и обнаружил, что за ночь было проверено **40 гипотез**. ![Терминал агента: Pursuing goal (1d 1h 41m)](https://anvilarth.github.io/img/sankalp_17pm.webp) _Так это выглядит со стороны: строчка в терминале и счётчик, который тикает, пока ты спишь. Кадр не мой — из [разбора sankalp](https://sankalp.bearblog.dev/autoresearch/), про него следующая глава._ Не все сорок оказались полезными — часть была откровенным мусором, часть повторяла то, что я уже пробовал. Но это было неважно. Всю ночь я не делал руками ничего — а на выходе получил больше проверенных вариантов, чем сделал бы сам за неделю с ноутбуком. И стало ясно: дело было не в том, чтобы лучше управлять агентом. Дело было в том, кто оценивает результат. Сработало это один раз, ночью, на одной конкретной задаче. Осталось понять, что именно сработало — и, как ни странно, помогает это понять не собственный опыт, а то, что кто-то другой наткнулся почти на то же самое чуть раньше меня. ## Глава 2. У этого уже было имя То, что случилось той ночью, не было изобретением. У этого уже было имя: autoresearch. И пара громких прецедентов до меня. Сначала я наткнулся на [пост sankalp про QR-разложение](https://sankalp.bearblog.dev/autoresearch/). Задача классическая: разложить матрицу быстро. Он не трогал кернелы руками — натравил на задачу Codex как автономного агента. Тот держал beam из кандидатов, плодил суб-агентов под профилировку, матан и генерацию идей, сам отбраковывал тупые ветки по ходу. Итог — **232x speedup** на batched QR и 12-е место из 183 участников, при том что у автора не было профессионального GPU-опыта. ![Схема beam-поиска: ветки кандидатов, точки слияния, лучшая ветка](https://anvilarth.github.io/img/sankalp_10pm.webp) _Beam из кандидатов у sankalp: ветки живут параллельно, слабые гаснут, сильные сливаются. Никто не выбирает победителя руками — его выбирает критерий. Схема из [поста sankalp](https://sankalp.bearblog.dev/autoresearch/)._ И пока читал, дошло, что дело не в агенте и не в промпте. Раньше эксперт вкладывался в само решение: признаки собирал руками, эвристики придумывал из головы, кернелы писал нативно. Знать, как устроена задача, и было работой. Теперь это знание напрямую почти не нужно. Нужно уметь собрать цикл, который решение найдёт сам: бенчмарк, oracle, критерий остановки. Ты не пишешь кернел — ты строишь loop, который его перебирает. Тот же _bitter lesson_, только не про архитектуру модели, а про то, чем занят человек. Отсюда два прямых вывода. Первый: можно физически делать меньше работы. Та самая рутина тестирования гипотез, ради которой я ленился вечером, делегируется — агент гоняет её ночью, я сплю. Второй: работа делается лучше. Не «быстрее то же самое», а качественнее — oracle за то же время перебирает пространство решений шире, чем один человек, и находит варианты, до которых я бы сам просто не дошёл. И sankalp — не единственный. Karpathy собрал [autoresearch](https://github.com/karpathy/autoresearch): 630-строчный скрипт, автономный loop — агент читает training code, предлагает изменение, гоняет пятиминутный run, меряет улучшение, коммитит или откатывает, и по новой. За ночь — сотни экспериментов, ни одного клика. [AlphaEvolve](https://deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/) от DeepMind — та же идея в большем масштабе: Gemini плюс evolutionary framework плюс автоматические оценщики. Улучшили решения для 20% из 50 открытых математических задач, ускорили датацентры, чип-дизайн и AI-training у Google. ![График прогресса autoresearch: 83 эксперимента, 15 удержанных улучшений](https://anvilarth.github.io/img/karpathy_progress.png) _Одна ночь в репозитории Karpathy: 83 эксперимента, из них 15 удержанных улучшений. Серые точки — то, что loop отбраковал сам. Ровно та пропорция мусора, которую я утром увидел у себя в логе. График из [karpathy/autoresearch](https://github.com/karpathy/autoresearch)._ У Karpathy в README есть строчка, которая описывает этот сдвиг точнее, чем я формулировал сам: > You're not touching any of the Python files like you normally would as a researcher. Instead, you are programming the `program.md` Markdown files that provide context to the AI agents and set up your autonomous research org. Файл, который правит агент, и файл, который правит человек, — разные файлы. Вся работа исследователя переехала во второй. Разные названия, разные реализации, разный масштаб. Но один и тот же паттерн: экспертиза переезжает с «найти решение» на «построить систему, которая находит решение». И это пока рамка, не ответ. Чтобы понять, где у рамки края, мне помогли не чужие кейсы, а два своих. Один прошёл гладко — там почти не о чем рассказывать. Второй — тот, где пришлось спорить с собой: я был уверен, что знаю, как лучше. ## Глава 3. Я знал лучше Две недели я собирал фильтры для датасета руками. Задача простая на вид — решить, какие семплы выкинуть из обучения, какие оставить. На деле смесь инженерии и кураторства: читаешь статьи, смотришь на данные, формулируешь гипотезу, проверяешь, отбрасываешь, по новой. Я считал, что понимаю эту задачу. Вооружённый эго и знанием, как будет лучше, я две недели вкладывал интуицию и чтение статей в ручные фильтры. А потом я сделал ровно то, что делал в тот ленивый вечер — построил бенчмарк через Gemini как oracle и натравил на него VLM zero-shot. Просто чтобы посмотреть, побьёт ли он мою ручную работу. Никаких эвристик и подсказок, никакого моего участия. Первая попытка уже была лучше двух недель ручной работы. Вторая — ещё лучше. Это был не какой-то космический отрыв. Не «агент изобрёл фильтр, которого я не мог придумать». Просто результат был объективно лучше, а времени на него ушло — пара прогонов против двух недель моей жизни. И тут стало по-настоящему обидно — не потому что метод сработал, а потому что я две недели спорил с самим собой над решением, которое автоматика перебрала за пару попыток. До AlexNet признаки собирали руками. Инженеры придумывали эвристики, выковыривали фичи, настраивали пайплайны — и это была экспертиза, за которой стояли годы работы. Нейросети это автоматизировали: фичи стали учиться, а работа сместилась на архитектуру и данные. Теперь автоматизируется сама экспертиза исследователя. Не нужно идеально знать домен, чтобы принимать хорошие решения — достаточно построить oracle, который знает, что считать хорошим решением, и пустить агента. Это и есть bitter lesson на собственной шкуре. Две недели, которые побила одна ночь. Второй кейс дался куда спокойнее. Может, ставки были другие — а может, потому что после первого я перестал спорить с oracle. ## Глава 4. Не серебряная пуля Второй кейс был про ускорение обучения. Никакой внутренней драмы — просто задача, где frontier-модели дороги и могут быть оверкилом, но если направить их на оптимизацию времени инференса или обучения, они экономят деньги и время. Я оставил модель работать на два дня и получил **30% ускорение**. Дальше ускорять стало сильно сложнее, и смысла тратить на это время уже не было. Гладко, буднично, без истории про эго. Но даже на гладком кейсе — и тем более после — всплывали подводные камни, из-за которых метод не стоит воспринимать как серебряную пулю. Первый — loss parity. Я забыл добавить в oracle проверку, что loss не разъезжается. Агент честно сообщил «готово, ускорение есть» — а по факту loss везде был NaN. Модель формально стала быстрее, но перестала учиться. После того как я добавил loss parity check, проблема ушла. Банальная вещь, которую я тем не менее пропустил, потому что слишком доверил агенту решать, что считать успехом. Второй — недостаточно детальный oracle. Это всплыло позже, на задаче удаления текста с изображений. Я тестировал разные модели, но валидировал результат по полному изображению. Мелкие артефакты — остатки букв, полустёртый символ — я замечал глазами, а oracle нет. Он считал, что текст удалён, и решение, которое он называл лучшим, на деле оставляло за собой мусор. Как только я добавил в oracle кроп области, где текст должен был удалиться, — стал валидировать не картинку целиком, а именно факт удаления — лучшим решением оказался совсем другой пайплайн. Прежний победитель просто эксплуатировал слепоту oracle. ![График одной сессии на KernelBench-Mega: ускорение относительно референса против потраченных токенов](https://anvilarth.github.io/img/kernelbench_session.jpg) _Как выглядит правильно устроенная сессия — разбор одного прогона на [KernelBench-Mega](https://kernelbench.com/mega). Первые 64% времени вообще нет кода: замер базовой линии, микробенчи барьеров, вывод roofline. Первый работающий кернел появляется на отметке 224k токенов. И ближе к концу — точка, ради которой я это показываю: «finer split-K regresses → measured, reverted». Гипотеза не сработала, замер это поймал, откат._ Тот же прогон описан автором бенчмарка одной фразой: > It spent 64% of the session in silence timing the baseline, microbenchmarking grid barriers, deriving a ~29x bytes/token roofline. The one regression it tried (finer split-K) it measured and reverted instead of rationalizing. Measured and reverted instead of rationalizing — это и есть работающий oracle. Мой NaN получился ровно потому, что мерять было нечем, и агенту оставалось только рационализировать. Общий урок простой и неприятный: конструирование autoresearch — тонкая задача, где именно на проектирование oracle и loop'а нужно тратить время. Время уходит не на рулёжку агентом и не на поиск магического промпта, а на то, чтобы честно и подробно описать, что считать хорошим решением. Воспринимать это как автоматическую серебряную пулю — значит рано или поздно получить NaN на месте loss'а. И даже когда oracle сконструирован правильно и loop работает гладко, остаётся ещё одна переменная, которую нельзя убрать проектированием — сама модель внутри loop'а. Я гонял весь этот выверенный процесс на разных моделях, и оказалось, что при одном и том же oracle и одной и той же задаче результат зависел не только от того, что я построил, но и от того, кому я это доверил. ## Глава 5. Не мощность, а интуиция На ускорении кода у меня было три модели: Fable 5, Claude Opus 4.8 и GPT-5.5. Одна и та же задача, один и тот же oracle. Fable 5 — я просто описал задачу, и всё сработало. Агент выдавал гипотезы, которые имели смысл, проверял их, отбраковывал тупые. Мне оставалось только кивать на разумные шаги. Claude Opus 4.8 предложил решение, которое вызывало OOM, и на этом как бы сдался — не откорректировал сам, не попробовал другой путь. GPT-5.5 в итоге справился, но мне приходилось явно подсказывать, куда смотреть, иначе он топтался. Разница не в способности выполнять шаги. Все три умеют читать код, писать код, запускать, мерять. Разница в domain-интуиции — в том, насколько глубоко в модель «зашито» понимание, что в этой задаче вообще имеет смысл пробовать. У Fable 5 это понимание было. У двух других — нет. Я компенсирую нехватку интуиции у моделей тем, что закидываю в промпт контекст — блоги, гайды, чужие кейсы. Но это ограничение, и неприятное: чтобы знать, чем помочь модели, я должен сам заранее хорошо разбираться в теме. То есть ровно та экспертиза, которую я думал автоматизировать, мне приходится держать в голове, чтобы вытаскивать её кусками и подсовывать агенту. И это не только моё наблюдение. Тот же паттерн виден в независимых бенчмарках. [KernelBench-Mega](https://kernelbench.com/mega) — бенчмарк на whole-block megakernels, где нужно слить целый блок модели в один кернел, а не оптимизировать отдельные операции. Все предыдущие модели «выигрывали» на нём только за счёт multi-kernel Triton pipeline — технически проходит тест, но честного слияния нет, это обход. И только Fable 5, [по твиту автора бенчмарка](https://x.com/elliotarledge/status/2072814573753975266), написал первый настоящий megakernel. Прямая параллель с моим подводным камнем из четвёртой главы: плохо спроектированный oracle позволял жухлерство, пока кто-то не сделал честно. [AutoKaggle](https://arxiv.org/abs/2410.20424) — тот же паттерн вне GPU: multi-agent фреймворк для Kaggle-соревнований, oracle там — это само соревнование и тесты, и на восьми соревнованиях получилось 0.85 validation submission rate и 0.82 comprehensive score, на уровне человека. ![Твит Elliot Arledge про первый настоящий megakernel](https://anvilarth.github.io/img/tweet_arledge.png) _Автор бенчмарка перечисляет, кто и на сколько «выигрывал» до этого: Opus 4.8 — 14.4x, GLM-5.2 — 11.1x, GPT-5.5 — 4.3x, Sonnet 5 — 4.0x. Все — через multi-kernel Triton pipeline, который не проходит authenticity gate. [Твит целиком](https://x.com/elliotarledge/status/2072814573753975266)._ Получается, лучший инструмент — это не самая мощная модель, а модель с правильной интуицией под конкретную задачу. И чем детальнее сконструирован oracle, тем яснее видно, у кого эта интуиция есть, а у кого её приходится компенсировать контекстом. Дальше это движется к recursive loops — циклам, где агент улучшает не веса модели, а саму систему: код, пайплайн, артефакт. Каждое улучшение облегчает следующее. Надеюсь, скоро появятся инструменты, которыми можно будет просто пользоваться. Пока приходится каждый раз выбирать, какой loop под какую задачу, и работает это точечно. Что это тогда говорит про мою собственную роль, если интуиция — единственное, что пока ещё моё, — становится товаром, который можно закинуть в промпт? ## Глава 6. Что остаётся На вопрос «заменит ли AI нас» у меня теперь скучный ответ: для части работы это уже случилось. Стажёры и джуны раньше рисовали графики и проверяли простые гипотезы — сейчас это целиком делает агент. Но это про джунов. Со мной сложнее. Три года назад я три дня копался в исходниках PyTorch Lightning, чтобы разобраться, как там устроен checkpointing. Три дня — и я знал это на уровне, на котором мог объяснить кому угодно и починить что угодно. Сейчас то же самое агент делает за час. Освобождает от рутины — да. Но забирает кое-что взамен. То чувство авторства, которое раньше давало ручное ковыряние в исходниках — «я это понял, я это сделал» — его больше нет. По факту я ничего не сделал, это всё агенты. Это похоже на переход от роли сеньора к роли менеджера. История ровно та же, что и в обычной карьере: становишься тимлидом — и руками уже почти ничего не делаешь. Раньше сам сидел и решал, теперь смотришь, как решают другие, и следишь, чтобы решали в нужную сторону. На выходе получается больше, чем сделал бы сам. Но своего, руками сделанного, в этом больше нет. Позиция менее дофаминовая: прежнего драйва от собственноручного решения не будет. И я думал, что на этом можно остановиться. Менеджер-то всё равно нужен. Не потому, что он умнее тех, кем управляет, а потому что знает, куда направить — у него domain-интуиция. Это и был мой вывод в прошлой главе: разница не в мощности, а в интуиции. Там, где у модели нет интуиции, я компенсирую закидыванием контекста в промпт. Интуиция была последним, что оставалось моим. А потом за пару месяцев набралось несколько историй, которые в эту картину не укладываются. Jacobian Conjecture — классическая гипотеза алгебры. Если у многочленного отображения якобиан ненулевая константа, то отображение обратимо. Недавно — контрпример в размерности три. И нашла его не группа математиков. Нашла Fable 5, пока я смотрел финал чемпионата мира по футболу. Теренс Тао, лауреат Филдсовской премии, написал [пост](https://terrytao.wordpress.com/2026/07/21/a-digestion-of-the-jacobian-conjecture-counterexample/), чтобы это переварить. Он объясняет ретроспективно, геометрически, то, что модель нашла сама. И вот что меня поразило: Тао садится и считает, насколько это невероятно. Многочлен степени семь. У якобиана априори до **1329** ненулевых коэффициентов — при **360** степенях свободы общего многочленного отображения той же степени. 1329 уравнений, которые должны занулиться одновременно, в пространстве из 360 переменных. Брутфорсом такое не находится — шанс ткнуть в нужную точку нулевой. Значит, Fable не перебирала. Она видела структуру. ![Абзац из поста Теренса Тао с подсчётом 1329 уравнений против 360 степеней свободы](https://anvilarth.github.io/img/tao_1329.png) _Тот самый абзац у [Тао](https://terrytao.wordpress.com/2026/07/21/a-digestion-of-the-jacobian-conjecture-counterexample/). Последняя фраза — приговор перебору: «finding such a polynomial looks highly unlikely to be located by brute force»._ В конце поста стоит дисклеймер, который лет пять назад в тексте филдсовского лауреата смотрелся бы дико: > I used an AI chatbot to discuss various aspects of this problem and to confirm several of the calculations made here. Увидеть, где лежит решение, в пространстве, где наивный перебор бесполезен. Это ровно то, что я считал своим. В университете мне рассказывали, что есть разные уровни запоминания материала. Первая стадия — ты учишь термины: «якобиан», «многочлен», «обратимое отображение». Знаешь, что они есть, но не понимаешь. Вторая — ты знаешь весь материал: определения, теоремы, доказательства. Но не можешь применить — на экзамене заваливаешь задачу, потому что она чуть отличается от шаблона. И есть третья стадия, которая и есть знание материала — интуиция. Не просто знать, что факт есть, а видеть, когда его применить, и узнавать ту же структуру в новой ситуации. Модели проделали тот же путь. Сначала стохастические попугаи — угадывали следующее слово. Потом выучили материал, знают факты, но ломаются на чём-то чуть новом. И вот теперь — третья стадия. Видят структуру, находят решения, которых не было в обучающей выборке. Контрпример к гипотезе — это не воспроизведение того, что модель видела, это новая математика. И схитрить тут негде: контрпример либо работает, либо нет, плохо спроектированного oracle, который можно обойти, в чистой математике не существует. Ладно, допустим, интуиция у модели есть. Но менеджер, который направляет, всё равно нужен — на этом я и держался. Второй случай ровно про это, и он мне ближе, потому что там видно, что делал человек. Математик Дмитрий Рыбин взял гипотезу Динница-Гарга-Гоманса — задача про сетевые потоки, висела с девяностых: можно ли дробное решение сделать неделимым, не подняв стоимость и не перегрузив ни одно ребро больше, чем на размер самой большой отправки. Контрпример нашёл [GPT-5.6 Pro](https://chatgpt.com/share/6a60b2eb-0b64-83ee-9c76-7931ca1de063): граф на семь узлов, три отправки, дробная стоимость 58, а любая неделимая маршрутизация — минимум 60. ![Контрпример к гипотезе Динница-Гарга-Гоманса: граф на семь узлов с потоками, стоимостями и требованиями](https://anvilarth.github.io/img/rybin_graph.jpg) _Весь контрпример целиком: семь узлов, один источник s, три получателя с требованиями 15, 10 и 15. Синие числа — поток, красные — стоимость ребра. Тридцать лет никто не мог такой граф ни построить, ни доказать, что его нет. [Картинка из твита Рыбина](https://x.com/DmitryRybin1/status/2079904005652893709)._ Интересна тут не модель, а промпты. Их было четыре, суммарно меньше шестидесяти слов: «construct a counterexample», «you should do a breakthrough», «it's enough of partial results, let's finish». Три захода подряд по полтора часа модель возвращала частичные результаты, и человек просто просил продолжать. Domain-интуиции в этих словах нет вообще. Есть направление и понимание, когда результат уже годится, — то есть ровно то, что я оставил себе как менеджеру. Оказалось, этой работы примерно на минуту. Тот же навык переносится и на мою работу. Мой личный труд — что он такое? Набор паттернов, которые я наработал за годы. «Видел эту ошибку — знаю, куда смотреть». «Эта архитектура ломается вот так — обходи вот так». «Эта гипотеза не сработает, потому что вот эта метрика врёт». Это всё паттерны. И их можно выучить. И сеть, которая нашла решение в пространстве из 1329 уравнений, — она выучит и мои паттерны. Вопрос лишь в цене, и цена эта каждый месяц дешевле, чем вчера. И вот что странно: кайф при этом никуда не делся. Я всё так же с утра открываю лог и залипаю в то, что там за ночь произошло, хотя руками не сделал ничего. Значит, кайф был не от «я решил это сам», а от того, что задача сдвинулась. А сдвигать её теперь получается сильно больше: бóльшую часть обучения моделей я тяну один, там, где раньше нужна была команда. Поэтому я не бросаю, а собираю следующий loop. Но на вопрос, в чём теперь моя ценность, ответа у меня нет. ### Источники - [sankalp — Auto-Research: QR Decomposition](https://sankalp.bearblog.dev/autoresearch/) — 232x speedup на batched QR (419 000 µs → 1805 µs) через Codex как автономного агента; 12-е место среди 183 участников, больше 1500 сабмитов за 14 дней. - [Andrej Karpathy — autoresearch](https://github.com/karpathy/autoresearch) — 630-строчный скрипт: training code → изменение → 5-минутный run → мера улучшения → commit/rollback → repeat. - [DeepMind — AlphaEvolve](https://deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/) — Gemini + evolutionary framework + automated evaluators; улучшены решения для 20% из 50 открытых математических задач. - [Terence Tao — A digestion of the Jacobian conjecture counterexample](https://terrytao.wordpress.com/2026/07/21/a-digestion-of-the-jacobian-conjecture-counterexample/) — контрпример в размерности 3; многочлен степени 7, до 1329 коэффициентов якобиана при 360 степенях свободы. - [KernelBench-Mega](https://kernelbench.com/mega) — независимый бенчмарк на whole-block megakernels (автор — Elliot Arledge). - [Elliot Arledge (твит)](https://x.com/elliotarledge/status/2072814573753975266) — первый настоящий megakernel на KernelBench-Mega, написанный Fable 5. - [AutoKaggle (arXiv:2410.20424)](https://arxiv.org/abs/2410.20424) — multi-agent фреймворк для Kaggle-соревнований; validation submission rate 0.85, comprehensive score 0.82 на 8 соревнованиях. - [Дмитрий Рыбин — контрпример к гипотезе Динница-Гарга-Гоманса](https://chatgpt.com/share/6a60b2eb-0b64-83ee-9c76-7931ca1de063) — четыре промпта, меньше 60 слов; три захода с частичными результатами, контрпример на четвёртом: граф на 7 узлов, три отправки, дробная стоимость 58 против минимум 60 у любой неделимой маршрутизации. --- # История о том, как штука для рисования треугольников стала самым важным чипом на планете _Андрей Филатов · 2026 · ~20 минут_ _Источник: https://anvilarth.github.io/gpu-story.html · Автор: Андрей Филатов (https://anvilarth.github.io/)_ Я работаю с GPU каждый день — запускаю модели, считаю инференс, жду пока обучится очередной эксперимент. Но долгое время для меня GPU был чёрным ящиком: закинул модель, нажал кнопку, ждёшь. А мне хочется понимать, что будет происходить дальше с миром AI — и я понял, что без понимания железа это невозможно. Я уже [пытался предсказать](https://anvilarth.github.io/realtime-video-calculator.html), когда видеогенерация станет реалтаймовой — получилось неплохо, и мне захотелось погрузиться вглубь: а что вообще стоит за GPU? А ведь решения в мире AI сейчас принимают те же люди, которые проектировали GPU. Если разобраться, **почему** они принимали именно такие решения на каждом этапе — можно научиться самому прогнозировать, куда всё движется. Ну и заодно построить ментальную модель GPU, которой мне так не хватало. > (а теперь SEO оптимизация) NVIDIA стоит дороже всех компаний мира. H100 продаётся за $30 000. И всё это — из-за процессора, который 25 лет назад рисовал треугольники в компьютерных играх. Как мы сюда попали — разбираемся с нуля. #NVIDIA #GPU #AI ## Как компьютер рисует картинку ### От текста к полигонам Для меня не секрет, что GPU были созданы для игр — первое что я узнал про видеокарты, это как раз NVIDIA 9800 GT, которые позволяли мне запускать любимые Assassin's Creed и Call of Duty. Но как понадобилось отдельное железо для отрисовки картинок? Первые компьютеры общались с человеком только текстом — зелёные буквы на чёрном экране. Zork (1977) — один из первых хитов: ты печатаешь `open mailbox` и получаешь текстовое описание того, что внутри. Такой DnD-симулятор квестов. Я их не люблю, и видимо не я один — потому что довольно быстро появились первые графические игры. _1977 · текст · 0 пикселей графики_ В начале 80-х появились [спрайты](https://ru.wikipedia.org/wiki/%D0%A1%D0%BF%D1%80%D0%B0%D0%B9%D1%82_(%D0%BA%D0%BE%D0%BC%D0%BF%D1%8C%D1%8E%D1%82%D0%B5%D1%80%D0%BD%D0%B0%D1%8F_%D0%B3%D1%80%D0%B0%D1%84%D0%B8%D0%BA%D0%B0)) — заранее нарисованные картинки, которые просто двигаются по экрану. Super Mario Bros (1985), Sonic (1991) — целые яркие миры, собранные из плоских кирпичиков. Это намного круче текстовых игр, но и плата — повышенные вычислительные аппетиты. Впрочем, со спрайтами терпимо: компьютер просто копирует готовые блоки в нужные позиции, 60 раз в секунду. Никаких вычислений «как это должно выглядеть» — только «куда положить картинку». _1985 · спрайты двигаются по экрану — копирование готовых блоков_ Но плоские картинки не передают глубину. Wolfenstein 3D (1992) и Doom (1993) красиво это обошли — мир хранился как плоская карта, а движок рисовал столбцы стен по расстоянию. Гениальный трюк, но в Doom нельзя даже посмотреть вверх — третьего измерения буквально нет. Настоящий прорыв — Quake (1996): полностью трёхмерный мир из полигонов. Свободная камера, текстуры, освещение — и всё собрано из сотен треугольников, каждый из которых нужно трансформировать, раскрасить и отрисовать. 30 раз в секунду. _1996 · 3D полигоны · трансформация + освещение каждого пикселя_ ### Один кадр Quake Давайте заглянем под капот. Quake на экране: игрок в коридоре, стены с текстурами, факел, в дверном проёме — враг. Камера может повернуться куда угодно. Всё это нужно превратить в картинку — 640 на 480 пикселей, 30 раз в секунду. У компьютера 33 миллисекунды на один кадр. Что происходит за эти 33 миллисекунды? Берём все треугольники сцены, пересчитываем координаты из 3D в плоские экранные. Определяем, какие пиксели попадают внутрь каждого треугольника. И для каждого из 307 200 пикселей считаем цвет: читаем текстуру, считаем освещение, проверяем глубину. Это дофига математических операций — возьми одно число, умножь на второе, сложи с третьим, и так для каждого пикселя. Анимация — как рисуется один кадр _640×480 · 30 fps · ~300M операций/сек. CPU рисует каждый пиксель последовательно._ Посчитаем. 640 × 480 = 307 200 пикселей. На каждый — 20–50 операций. На один кадр выходит 10–15 миллионов операций. Умножаем на 30 fps — 300–450 миллионов операций каждую секунду. Pentium 200 МГц — топовый процессор 1996 года — выдавал около 200 миллионов инструкций в секунду. Уже на грани. Но в реальности всё ещё хуже: каждый пиксель лезет в текстуру — изображение, натянутое на треугольник. Адрес непредсказуем, в кэше данных почти наверняка нет. Кэш-промах — процессор стоит и ждёт 50–100 наносекунд, пока память ответит. Десятки потерянных тактов. На каждый пиксель. Реальная производительность падает в 2–3 раза. Прям мега грустно. CPU явно не справляется — нужно отдельное железо для графики. ### Зашито в кремний Рынок это понял быстро. В середине 90-х появились первые графические ускорители: 3dfx Voodoo, ATI Rage, S3 ViRGE. Идея простая — вынести графические операции с CPU на отдельный чип. Но транзисторов мало, жили бедно, поэтому операции напрямую впаивали в кремний. Voodoo умел растеризировать и накладывать текстуры — и только это. Трансформация вершин? Всё ещё на CPU. ATI Rage умел и то, и другое, но только с одним источником света. S3 ViRGE вообще оказался медленнее программного рендеринга. Каждый чип — фиксированный набор функций. Хочешь эффект, которого в чипе нет — мечтай. Для разработчика — кошмар. Каждая карта говорит на своём языке: свои регистры, свои команды. Написал игру под 3dfx — на ATI она не работает. Вышла новая карта — переписывай всё заново. [OpenGL](https://ru.wikipedia.org/wiki/OpenGL) (1992) и [DirectX](https://ru.wikipedia.org/wiki/DirectX) (1995, эхх, сколько его приходилось ставить) частично решили проблему совместимости: разработчик вызывает стандартную функцию вроде `DrawTriangle`, а драйвер переводит её в команды конкретного чипа. Один код → любое железо. Но проблема fixed-function никуда не делась. Каждый ускоритель умел только то, что в него зашили. Хочешь новый эффект — жди следующее поколение чипов. И ни один из них не покрывал весь конвейер целиком: часть работы всё равно падала на CPU. Нужен был другой подход. ## Проектируем GPU с нуля Ок, хотим спроектировать ускоритель — надо сначала понять, что именно ускоряем. Рендеринг кадра — это конвейер. Сначала каждую вершину треугольника пересчитываем из 3D в экранные координаты — матричное умножение 4×4. Потом определяем, какие пиксели попадают внутрь треугольника — растеризация. И для каждого пикселя считаем цвет: текстура, освещение, глубина. Интерактивное демо — путь одного треугольника _Три вершины задают треугольник в трёхмерном пространстве._ И вот ключевое наблюдение, из которого родился GPU: на каждом этапе — тысячи операций, и все они **независимы**. Один пиксель не ждёт результата другого. Одна вершина не зависит от соседней. **Формула одинаковая, данные разные, всё независимо.** Когда это слово встречается — стоит подумать о параллелизме. Осталось спроектировать процессор, который умеет именно это. ### Почему не CPU Внутри CPU — нажми на блок А как с параллелизмом у CPU? Спойлер — всё плохо. Большая часть транзисторов CPU — это не вычисления, а управление. Предсказание ветвлений, переупорядочивание инструкций, кэши. Это делает CPU невероятно умным для сложных последовательных задач — как гений-одиночка. Но это как заставить профессора математики считать на калькуляторе. Он справится, но это не его сильная сторона. Нам не нужно предсказывать ветвления — формула для каждого пикселя одинаковая. Не нужно переставлять инструкции — все потоки делают одно и то же. Не нужны огромные кэши — данные всё равно непредсказуемы. Чисто вычислительная часть CPU занимает ~20% площади — ALU, арифметико-логическое устройство. Единственный блок, который реально считает. Ну вы как умный человек подумаете — мало вычислений? Давайте сделаем больше вычислений. Забьём весь чип ALU и всё будет по кайфу. И будете правы. ### Шаг 1: Множим ALU Берём CPU, выкидываем всё «умное», оставляем только ALU. На освободившуюся площадь можно уместить не одно ALU с обвязкой, а десятки голых. Каждый проще, каждый медленнее — но их много, и они работают одновременно. Не один умный процессор — целая армия простых. ### Шаг 2: Одна команда на всех Окей, у нас десятки ALU на одном кристалле. Но сразу столкнёмся с проблемой — блин, а как ими управлять? Если каждому дать свой декодер команд, свой счётчик инструкций, свою управляющую логику — мы опять потратим кучу транзисторов на управление. Тут погорячились — вернулись к тому, от чего убежали. Но вспомним нашу задачу: все пиксели считаются одной и той же формулой. Цвет пикселя 0 — как цвет пикселя 31, просто входные данные разные. Значит, можно проще: **один** декодер команд на группу ALU. Читает инструкцию один раз — все ALU в группе выполняют её одновременно, каждый со своими данными. Это [SIMD](https://ru.wikipedia.org/wiki/SIMD) — Single Instruction, Multiple Data. Все делают одно и то же. В GPU NVIDIA такая группа — **warp**, 32 потока. Почему 32? Компромисс: больше группа — дешевле управление, но больше проблем когда потоки хотят делать разные вещи. Итого: один декодер, 32 ALU, одна инструкция — все считают параллельно. Проблему управления решили. Но есть нюанс. Что если в коде `if/else`? При рендеринге Quake проверяем, попадает ли пиксель в тень. Часть из 32 пикселей в тени, часть на свету — а декодер-то один, он не может выполнять разные инструкции одновременно. Решение элегантное: GPU выполняет **обе ветки** на всех 32 ALU, но с маскированием. Прогоняет ветку «в тени» — все 32 считают, но результаты пикселей на свету отбрасываются. Потом «на свету» — то же самое наоборот. Два такта вместо одного, часть работы впустую. Это **warp divergence**. Для графики терпимо — соседние пиксели обычно в одной ветке. ### Шаг 3: Прячем задержки У нас массив ALU, объединённых в warps. Они считают быстро. Но есть узкое место: память. Когда шейдер обращается к текстуре — а это происходит для каждого пикселя — данные нужно прочитать из видеопамяти. А видеопамять медленная по сравнению с процессором: запрос и ответ занимают 200–400 тактов. Всё это время 32 ALU в warp стоят и ждут. Ничего не считают. Чистый простой. (Кстати, проблемы с HBM у H100 взялись не из ниоткуда — перегонять данные с памяти на чип было дорого ещё тогда.) CPU решает это кэшами — держит данные поближе. Но мы же выкинули большие кэши, когда проектировали GPU. Как быть? Решение красивое. Вместо того чтобы прятать задержку — GPU **заполняет** её полезной работой. На одном процессорном блоке (SM — Streaming Multiprocessor) GPU держит в регистрах состояние не одного warp, а **десятков**. Переключение — за один такт, без накладных расходов. Warp A запросил текстуру и ждёт? GPU переключается на warp B. Потом C. Потом D. К моменту, когда данные для A приходят из памяти — очередь как раз до него дошла. Простои, чтобы от них избавиться — GPU просто делает другую работу, пока ждёт. Чем больше warps в полёте — тем лучше скрывается задержка. Задержка 200 тактов, один warp считается 60 — значит, нужно минимум 4 в очереди. В реальности GPU держит 32–64 warps на одном SM и почти никогда не простаивает. Важнейший принцип — запоминаем. > CPU прячет задержку кэшами — хранит данные рядом. GPU — переключением потоков: пока один ждёт, другой считает. Подробнее об архитектуре GPU, warps и latency hiding — в бесплатном курсе [Cornell Virtual Workshop: Understanding GPU Architecture](https://cvw.cac.cornell.edu/index). ### Мы только что создали GPU Оглянемся. За три шага мы спроектировали новый тип процессора. Выкинули из CPU всё лишнее для графики — предсказание ветвлений, перестановку инструкций, огромные кэши. Забили площадь десятками простых ALU. Объединили их в группы с общим декодером — SIMD. Проблему медленной памяти решили переключением между потоками. Это и есть GPU. Именно так рассуждали инженеры NVIDIA в конце 90-х. Складываем всё вместе — и в 1999 году появляется GeForce 256. Они сами назвали его «первым в мире GPU». 17 миллионов транзисторов, 120 МГц, 10 миллионов полигонов в секунду. Главное — аппаратный Transform & Lighting: трансформации вершин и расчёт освещения впервые целиком переехали с CPU на видеокарту. Один чип — весь конвейер рендеринга. ![NVIDIA GeForce 256 — первый в мире GPU, 1999](https://anvilarth.github.io/geforce256.png) _NVIDIA GeForce 256 (1999) · 17M транзисторов · 120 МГц · 10M полигонов/с · фото: Wikimedia Commons, public domain_ ## От шейдеров к CUDA ### Шейдеры: от зашитого к программируемому Мы нанесли удар по скорости вычислений — больше транзисторов, больше ALU, красота. Но была другая проблема: всё зашито в кремний. Хочешь другой эффект? Жди новый чип. Почему вообще зашивали? Бюджет транзисторов — 17 миллионов, каждый на счету. Вот конкретный пример. Берём один пиксель на стене в Quake рядом с факелом. Чтобы вычислить его цвет, чип проводит пиксель через цепочку зашитых операций: Проблема конкретная. Та формула, которую мы только что видели — она считает освещение стены. Свет падает, отражается, получается блик. Для кирпичной стены в Quake — отлично. Но вода отражает и преломляет свет совсем иначе. Кожа рассеивает свет под поверхностью. Металл — резкий блик, а мех — мягкий. Для каждого материала нужна своя формула. А в чипе зашита одна. И зашить все варианты в кремний невозможно — их слишком много. Но рост числа транзисторов помог: 17М (1999) → 57М (2001) → 107М (2002). Появился бюджет для эксперимента. И таким образом появились **шейдеры** — небольшие программы, которые пишет разработчик, а GPU исполняет для каждой вершины или пикселя. Разница принципиальная: в fixed-function транзисторы соединены проводами в определённом порядке — чип всегда выполняет ровно ту цепочку, которую впаяли на заводе. В шейдерной модели на чипе стоит универсальный ALU + память для инструкций. Перед рендерингом загружаешь туда свою программу — любую. Тот же ALU может посчитать и стену, и воду, и кожу — просто загрузи другую формулу. Заскейлились: _[схема] Рост программируемости: от 128 инструкций без ветвлений до полноценного процессора за 3 года._ **Shader Model 1.0** (2001, GeForce 3) — до 128 инструкций, нет ветвлений. По сути тот же fixed-function, только можно переставить шаги. Программируемый блок в 2-3 раза медленнее зашитого — гибкость стоит дорого. **Shader Model 2.0** (2002, Radeon 9700) — 256 инструкций, впервые `if/else`. Уже можно писать реальную логику, overhead меньше. **Shader Model 3.0** (2004, GeForce 6) — до 65 000 инструкций, циклы, динамическое ветвление. Полноценный процессор. Overhead стал настолько маленьким, что драйверы начали эмулировать fixed-function через шейдеры — зашитые блоки буквально потеряли смысл. Нафиг тогда в железо операции зашивать, если программно почти так же быстро и ещё миллион вещей умеет? Гибкость победила. > GPU из «калькулятора с кнопками» превратился в «программируемый калькулятор». _[схема] Unified architecture: вместо 8V+24P -- 128 универсальных блоков с динамическим распределением._ Следующий логичный шаг — все блоки на чипе стали одинаковыми. 128 универсальных блоков, планировщик сам раскидывает задачи. Утилизация близка к 100%. И тут инженеры NVIDIA осознали кое-что поинтереснее. Если все 128 блоков умеют любую математику — это уже не видеокарта. Это **массивно-параллельный вычислитель общего назначения**. Игровая оптимизация случайно создала суперкомпьютер на видеокарте. ### CUDA: как к этому вообще добраться Ок, у нас суперкомпьютер на видеокарте. Но попасть на него можно было только через графический API — OpenGL или DirectX. Хочешь посчитать уравнение? Данные закодируй как текстуру, вычисление оформи как шейдер, результат прочитай обратно как картинку. Учёный буквально притворяется что рисует картинку, чтобы GPU посчитал ему физику. Напоминает [Infinite Storage Glitch](https://github.com/KKarmugil/Infinite_Storage_Glitch) — проект, где ребята кодируют файлы в пиксели видео и заливают на YouTube как бесплатное облачное хранилище. Платформа для видео, но технически стримит любые байты. Решение появилось довольно быстро. В 2004 году Ian Buck в Стэнфорде сделал Brook — первый язык для GPU без графического API. Корявый, ограниченный — но доказал что идея работает. NVIDIA его наняла, и уже в 2007 вышла **CUDA**. Что изменилось принципиально? Вместо "закодируй данные как текстуру, притворись что рисуешь" — три простые идеи: **1. Просто напиши функцию** — помечаешь её `__global__` и говоришь "запусти на 10 000 потоков". Всё. GPU сам раскидает по ядрам. **2. Иерархия потоков** — threads собираются в blocks, blocks в grid. Ты думаешь о структуре задачи, а не о железе. А помните warps из Главы II? Вот как это связано: ты говоришь "блок из 256 потоков", а GPU внутри нарезает его на warps по 32 — те самые группы с общим декодером и SIMD. Программист управляет блоками, GPU управляет warps. **3. Потоки общаются** — внутри блока есть быстрая shared memory. Потоки могут обмениваться данными напрямую, не гоняя их через медленную глобальную память. По факту — пишешь нормальный C-код, помечаешь функцию `__global__`, компилируешь через `nvcc` и запускаешь. Те же 128 ядер считают твою задачу. Без текстур, без треугольников, без притворства. Просто математика. Железо не изменилось — изменилось то, как ты с ним разговариваешь. И учёные это оценили быстро, потому что паттерн тот же что в графике: одна формула на миллионы точек. GPU стал универсальным вычислителем. Но каждая новая задача обнажала слабые места — и требовала новых решений в железе. ## От вычислителя к AI-чипу Ок, у нас универсальный параллельный вычислитель с нормальным интерфейсом. Победа? Не совсем. Каждый раз когда учёные и инженеры начинали реально использовать GPU — они находили новую проблему. И каждое поколение архитектуры — это ответ на конкретную боль предыдущего. ### Fermi → Maxwell: GPU учится быть надёжным (2010–2014) Первая проблема оказалась неожиданной. Учёные запустили физическую симуляцию — а результат на GPU отличается от CPU. Не потому что алгоритм неправильный, а потому что в памяти случайно перевернулся бит. Для игр это незаметно — ну мигнул пиксель. Для науки — катастрофа, расчёт неверный. ### Pascal: одной карточки мало (2016) Модели росли быстрее памяти. Одна карточка — 12–16 ГБ, а нужно больше. Что делать? Поставить несколько рядом. Но они общались через PCIe — универсальный разъём на материнской плате, через который подключается вообще всё. Данные шли через CPU: GPU → CPU → GPU, всего ~16 ГБ/с. Pascal сделал прямое соединение GPU-GPU без крюка через CPU — NVLink, 160 ГБ/с, в 10 раз шире. Несколько карточек стали работать почти как одна. ### Volta: возврат специализации (2017) А вот тут легендарный разворот. Помните, в Главе III мы убрали фиксированные операции с GPU ради гибкости? А теперь NVIDIA возвращает их обратно — но уже для другой задачи. ML — это по сути умножение матриц, миллионы раз подряд. Обычное CUDA-ядро делает одну операцию за такт. Целая матрица 4×4 — 64 такта. Расточительно. ### Ampere: два мира на одном чипе (2020) Tensor Cores работали отлично, но мир изменился — появились модели на 175 миллиардов параметров (GPT-3). И выяснилось что тренировка и inference — совсем разные задачи. Тренировка хочет максимум вычислений, inference хочет минимум задержки. Один чип должен уметь и то и другое. Ampere A100 решил это тремя трюками. **TF32** — считает быстро как FP16, но точно почти как FP32, и главное без изменения кода. **Sparsity 2:4** — если в матрице половина нулей, GPU пропускает их аппаратно, 2× ускорение бесплатно. **MIG** — один GPU можно порезать на 7 изолированных кусков, на каждом своя задача. Первый чип, который проектировали одновременно для training и inference. ### Hopper: эра больших моделей (2022) LLM — GPT, LLaMA — оказались прожорливыми по-новому. Разные слои модели нуждаются в разной точности, но FP16 везде — расточительно: тратим память и bandwidth на точность которая не нужна. ### Blackwell: чип не может расти вечно (2024) Один монолитный чип упёрся в физику — чем больше кристалл, тем больше дефектов при производстве. Плюс 700W на чип — дата-центр из тысяч таких потребляет как небольшой город. ### Что дальше: Feynman (2028) NVIDIA уже показала роадмап. Следующий большой шаг — Feynman (2028). Горизонтально расти некуда — растём вертикально. ### Вся эволюция одним взглядом ## От картиночек до почти AGI Мы прошли путь от штуки, которая рисует треугольники в Quake, до чипов, на которых тренируют модели с сотнями миллиардов параметров. GPU начинался как специализированный ускоритель для одной задачи, стал универсальным вычислителем — а теперь снова специализируется, только уже под матричные умножения. И именно эта новая специализация породила вопрос: а может, GPU — не единственный ответ? ### Проблема inference Когда LLM генерирует текст — один токен за раз. На каждый токен модель читает все свои веса из памяти: миллиарды параметров. Вычислений на прочитанный байт — крошечное количество. Для полной утилизации H100 нужно ~300 операций на каждый прочитанный байт, а реально получается ~1. Тысячи ядер и Tensor Cores простаивают, ожидая данные из HBM. Утилизация 5–15%. Купил карточку за $30K, а используешь на 10% — обидно. > Главная проблема: **как не ждать память**. TPU, Groq, Cerebras — все решают именно это, но по-разному. В GPU тысячи ядер, и каждое на каждом шаге бегает в общую память за данными. Посчитал — записал обратно. Посчитал — записал. Универсально, но ядра постоянно ждут пока память ответит. ### Google TPU: данные текут, а не стоят TPU устроен принципиально иначе. Давай пошагово на примере. В GPU каждое ядро на каждом шаге лезет в память: забрал число, умножил, записал обратно, забрал следующее. В TPU так: **Шаг 1.** Загружаем веса из памяти в решётку — один раз. Каждый элемент решётки запоминает свой вес. Элемент [0,0] запомнил W₁, элемент [0,1] запомнил W₂, и так далее. Веса "застыли" в решётке. **Шаг 2.** Пускаем входные данные X₁ слева в первый элемент. Он умножает X₁ на свой W₁ и передаёт результат соседу справа. Не в память — напрямую соседу. **Шаг 3.** Сосед получает результат, прибавляет свой W₂×X₂, передаёт дальше. Следующий — то же самое. На выходе справа — готовый результат строки матрицы. **Шаг 4.** Пока первая строка "проходит" через решётку, сзади уже запускается вторая. Элементы не простаивают ни одного такта. Итого: GPU обращается к памяти **на каждую операцию**. TPU — **один раз на всю матрицу**. Разница в количестве обращений огромная, и именно поэтому для матричных умножений TPU быстрее. Главная ставка Google — масштаб: TPU Pod из тысяч чипов, соединённых быстрой сетью. Потенциал: если ты в экосистеме Google Cloud — может быть дешевле и быстрее GPU. Но ты привязан к Google, своё не поставишь. ### Groq LPU: убрали память вообще Groq решил проблему ещё радикальнее. Помните из Главы II — HBM это внешняя память, стоит рядом с чипом но физически отдельно. Каждое обращение — сотни тактов задержки. SRAM — память прямо внутри чипа, на том же кристалле, доступ за 1-2 такта. Groq подумал: если при inference мы просто читаем веса по порядку и модель влезает в SRAM — зачем вообще ходить во внешнюю память? Убрали HBM, всё на чипе. Задержка исчезла, скорость inference детерминистическая. Но ограничение очевидное: SRAM конечный, большие модели не влезают. ### Cerebras WSE: чип размером с пластину Cerebras зашёл с третьей стороны: если проблема в том что данные далеко от вычислений — сделаем чип гигантским, чтобы всё влезло. Один кристалл размером с целую кремниевую пластину — 850 000 ядер, 44 ГБ SRAM. Зачем нарезать пластину на маленькие чипы и потом соединять обратно, если можно оставить как есть? Потенциал: невероятная плотность. Но производство сложное и экосистема маленькая. ### Что из этого следует В начале этого поста GPU был для меня чёрным ящиком. Закинул модель, нажал кнопку, ждёшь. Теперь — нет. Мы прошли от рисования треугольников в Quake до чипов, на которых тренируют модели, приближающие нас к AGI. И за каждым решением — не магия, а конкретная инженерная проблема и конкретный компромисс. Для меня главный вывод такой: когда разбираешь что-то до базовых принципов — перестаёшь бояться сложного. GPU казался невероятно запутанной штукой, а оказался набором остроумных решений под ограничения физики. Параллелизм — потому что пиксели независимы. SIMD — потому что формула одинаковая. Latency hiding — потому что память медленная. Tensor Cores — потому что ML это матрицы. Любая технология так устроена — разбери до первопринципов и станет понятно. И это навык, который переносится на всё остальное. Но знать архитектуру — половина дела. Вторая — уметь писать код, который её учитывает. Как самому написать CUDA kernel. Какие библиотеки и решения существуют. Как понять, что твой код использует GPU на 10%, и что с этим делать. Об этом — в следующем посте. ### Источники - [CMU 15-462 — How a GPU Works](https://www.cs.cmu.edu/afs/cs/academic/class/15462-f11/www/lec_slides/lec19.pdf) - [Fabian Giesen — A trip through the Graphics Pipeline](https://fgiesen.wordpress.com/2011/07/09/a-trip-through-the-graphics-pipeline-2011-index/) - [CMU 15-418 — GPU Architecture](http://15418.courses.cs.cmu.edu/spring2015/lecture/gpuarch/slide_003) - [NVIDIA Blog — 25th Anniversary of GeForce 256](https://blogs.nvidia.com/blog/first-gpu-gaming-ai/) - [Wikipedia — CUDA](https://en.wikipedia.org/wiki/CUDA) - [Yahoo Finance — Going all-in with Nvidia](https://finance.yahoo.com/news/going-all-in-with-nvidia-how-jensen-huangs-high-stakes-bets-paid-off-113053891.html) - [The Chip Letter — Nvidia's Embarrassingly Parallel Success](https://thechipletter.substack.com/p/nvidias-embarrassingly-parallel-success) - [Acquired Podcast — Nvidia Part I](https://www.acquired.fm/episodes/nvidia-the-gpu-company-1993-2006) - [CloudFleet — NVIDIA GPU Architectures](https://cloudfleet.ai/blog/cloud-native-how-to/2023-03-comparison-of-different-nvidia-gpu-rchitectures/) - [Google — In-Datacenter Performance Analysis of a Tensor Processing Unit (TPU paper)](https://arxiv.org/abs/1704.04760) - [Pope et al. — Efficiently Scaling Transformer Inference](https://arxiv.org/abs/2211.05102) - [NVIDIA — Volta Architecture Whitepaper (Tensor Cores)](https://images.nvidia.com/content/volta-architecture/pdf/volta-architecture-whitepaper.pdf) - [NVIDIA — Hopper Architecture Whitepaper (H100, Transformer Engine)](https://resources.nvidia.com/en-us-grace-hopper/gtc24-whitepaper-hopper) - [NVIDIA — Blackwell Architecture Whitepaper (B200, 208B transistors)](https://resources.nvidia.com/en-us-blackwell-architecture) - [NVIDIA — NVLink & NVSwitch](https://www.nvidia.com/en-us/data-center/nvlink/) - [Wikipedia — Brook (Ian Buck, Stanford GPGPU)](https://en.wikipedia.org/wiki/Brook_(programming_language)) - [Wikipedia — Shader History (Shader Model evolution)](https://en.wikipedia.org/wiki/Shader#History) - [Groq — LPU Architecture (SRAM-only, deterministic inference)](https://groq.com/technology/) - [Cerebras — Wafer-Scale Engine (850K cores, 44GB SRAM)](https://www.cerebras.net/chip/) - [Wikipedia — Systolic Array](https://en.wikipedia.org/wiki/Systolic_array) --- # How a thing that drew triangles became the most important chip on the planet _Andrey Filatov · 2026 · ~20 min read_ _Source: https://anvilarth.github.io/gpu-story.html · Author: Andrei Filatov (https://anvilarth.github.io/)_ I work with GPUs every day — running models, doing inference, waiting for yet another experiment to train. But for a long time the GPU was a black box to me: throw in a model, press the button, wait. I want to understand where the AI world is heading — and I realized that's impossible without understanding the hardware. I've already [tried to predict](https://anvilarth.github.io/realtime-video-calculator.html) when video generation would go real-time — it turned out pretty well, and I wanted to go deeper: what actually stands behind the GPU? The decisions in the AI world today are made by the same people who designed GPUs. If you figure out **why** they made those exact decisions at each step — you can learn to predict where everything is heading yourself. And build that mental model of the GPU I've been missing. > (and now some SEO) NVIDIA is worth more than any company on the planet. An H100 sells for $30,000. All because of a chip that 25 years ago was drawing triangles in video games. How did we get here — let's figure it out from scratch. #NVIDIA #GPU #AI ## How a computer draws a picture ### From text to polygons It's no secret to me that GPUs were made for gaming — my first encounter with graphics cards was the NVIDIA 9800 GT that let me run my favorite Assassin's Creed and Call of Duty. But why did we need separate hardware just to draw pictures? Early computers communicated with humans through text only — green letters on a black screen. Zork (1977) was one of the first hits: you type `open mailbox` and get a text description of what's inside. A DnD-style quest simulator. I'm not a fan, and apparently I'm not alone — because graphical games appeared pretty quickly after. _1977 · text · 0 pixels of graphics_ In the early 80s [sprites](https://en.wikipedia.org/wiki/Sprite_(computer_graphics)) appeared — pre-drawn images that simply move around the screen. Super Mario Bros (1985), Sonic (1991) — entire vibrant worlds built from flat tiles. Way cooler than text games, but the price was higher computational appetite. Though with sprites it was manageable: the computer just copies pre-made blocks to the right positions, 60 times per second. No calculations for "how should this look" — just "where to place the image". _1985 · sprites moving on screen — copying pre-made blocks_ But flat images can't convey depth. Wolfenstein 3D (1992) and Doom (1993) worked around this elegantly — the world was stored as a flat map, and the engine drew wall columns based on distance. A brilliant trick, but in Doom you can't even look up — the third dimension literally doesn't exist. The real breakthrough was Quake (1996): a fully 3D world made of polygons. Free camera, textures, lighting — all built from hundreds of triangles, each one needing to be transformed, colored, and rendered. 30 times per second. _1996 · 3D polygons · transform + light every pixel_ ### One frame of Quake Let's look under the hood. Quake on screen: a player in a corridor, textured walls, a torch, an enemy in the doorway. The camera can turn anywhere. All of this needs to become an image — 640 by 480 pixels, 30 times per second. The computer has 33 milliseconds per frame. What happens in those 33 milliseconds? We take all the triangles in the scene, convert coordinates from 3D to flat screen space. Figure out which pixels fall inside each triangle. And for each of the 307,200 pixels we compute the color: read the texture, calculate lighting, check depth. That's a ton of math operations — take one number, multiply by another, add a third, and repeat for every single pixel. Animation — how one frame is drawn _640×480 · 30 fps · ~300M ops/sec. The CPU draws each pixel sequentially._ Let's do the math. 640 × 480 = 307,200 pixels. Each one takes 20–50 operations. One frame comes out to 10–15 million operations. Multiply by 30 fps — 300–450 million operations every second. The Pentium at 200 MHz — the top CPU of 1996 — could do about 200 million instructions per second. Already at the limit. But in reality it's even worse: each pixel needs a texture lookup — an image mapped onto a triangle. The address is unpredictable, almost certainly not in the data cache. A cache miss — the processor just sits there waiting 50–100 nanoseconds for memory to respond. Dozens of wasted cycles. Per pixel. Real-world performance drops 2–3x. Truly depressing. The CPU clearly can't handle this — we need dedicated hardware for graphics. ### Baked into silicon The market figured this out fast. In the mid-90s the first graphics accelerators appeared: 3dfx Voodoo, ATI Rage, S3 ViRGE. The idea was simple — offload graphics operations from the CPU to a dedicated chip. But transistors were scarce, budgets were tight, so operations were hardwired directly into silicon. The Voodoo could rasterize and apply textures — and that's it. Vertex transforms? Still on the CPU. ATI Rage could do both, but only with a single light source. The S3 ViRGE actually turned out slower than software rendering. Each chip was a fixed set of functions. Want an effect that's not in the chip — keep dreaming. For developers — a nightmare. Every card speaks its own language: its own registers, its own commands. Wrote a game for 3dfx — it doesn't work on ATI. New card comes out — rewrite everything from scratch. [OpenGL](https://en.wikipedia.org/wiki/OpenGL) (1992) and [DirectX](https://en.wikipedia.org/wiki/DirectX) (1995, man, how many times I had to install it) partially solved the compatibility problem: the developer calls a standard function like `DrawTriangle`, and the driver translates it into commands for the specific chip. One codebase → any hardware. But the fixed-function problem didn't go anywhere. Each accelerator could only do what was hardwired into it. Want a new effect — wait for the next generation of chips. And none of them covered the entire pipeline: part of the work still fell on the CPU. A different approach was needed. ## Designing a GPU from scratch Ok, we want to design an accelerator — first we need to understand what exactly we're accelerating. Rendering a frame is a pipeline. First, we transform each triangle vertex from 3D into screen coordinates — a 4×4 matrix multiplication. Then we figure out which pixels fall inside the triangle — rasterization. And for each pixel we compute the color: texture, lighting, depth. Interactive demo — one triangle's journey _Three vertices define a triangle in 3D space._ And here's the key observation that gave birth to the GPU: at every stage there are thousands of operations, and they're all **independent**. One pixel doesn't wait for another's result. One vertex doesn't depend on its neighbor. **Same formula, different data, everything is independent.** When that word comes up — it's time to think about parallelism. All that's left is to design a processor that does exactly this. ### Why not a CPU Inside a CPU — click a block So how does the CPU handle parallelism? Spoiler — it's bad. Most of the CPU's transistors aren't for computing, they're for control. Branch prediction, instruction reordering, caches. This makes the CPU incredibly smart for complex sequential tasks — like a lone genius. But it's like making a math professor punch numbers on a calculator. He'll manage, but it's not his strong suit. We don't need branch prediction — the formula is the same for every pixel. No need for instruction reordering — all threads do the same thing. No need for huge caches — the data is unpredictable anyway. The pure compute portion of a CPU takes up ~20% of die area — the ALU, the arithmetic logic unit. The only block that actually computes. Being a smart person, you'd think — not enough compute? Let's just add more compute. Pack the entire chip with ALUs and we're golden. And you'd be right. ### Step 1: Multiply the ALUs Take a CPU, throw out all the "smart" stuff, keep only the ALU. On the freed-up area you can fit not one ALU with its overhead, but dozens of bare ones. Each is simpler, each is slower — but there are many of them, and they all work simultaneously. Not one smart processor — an entire army of simple ones. ### Step 2: One instruction for all Ok, we've got dozens of ALUs on a single die. But we immediately hit a problem — wait, how do we control them all? If each one gets its own instruction decoder, its own program counter, its own control logic — we'll burn a ton of transistors on control again. Got ahead of ourselves — back to square one. But remember our task: every pixel is computed with the same formula. Pixel 0's color — same as pixel 31's, just different input data. So we can do it simpler: **one** instruction decoder per group of ALUs. It reads the instruction once — all ALUs in the group execute it simultaneously, each with its own data. This is [SIMD](https://en.wikipedia.org/wiki/Single_instruction,_multiple_data) — Single Instruction, Multiple Data. Everyone does the same thing. In NVIDIA GPUs this group is called a **warp** — 32 threads. Why 32? It's a tradeoff: larger group — cheaper control, but more issues when threads want to do different things. Bottom line: one decoder, 32 ALUs, one instruction — all computing in parallel. Control problem solved. But there's a catch. What if the code has an `if/else`? When rendering Quake we check whether a pixel falls in shadow. Some of the 32 pixels are in shadow, some in light — but there's only one decoder, it can't execute different instructions at the same time. The solution is elegant: the GPU executes **both branches** on all 32 ALUs, but with masking. It runs the "in shadow" branch — all 32 compute, but results for lit pixels are discarded. Then "in light" — same thing in reverse. Two cycles instead of one, some work wasted. This is **warp divergence**. For graphics it's tolerable — neighboring pixels usually take the same branch. ### Step 3: Hiding latency We've got an array of ALUs grouped into warps. They compute fast. But there's a bottleneck: memory. When a shader reads a texture — and this happens for every pixel — it needs to fetch data from video memory. And video memory is slow compared to the processor: a request-response takes 200–400 cycles. During all that time, 32 ALUs in the warp just sit there waiting. Computing nothing. Pure idling. (By the way, the HBM issues with H100 didn't come out of nowhere — moving data from memory to chip was expensive even back then.) CPUs solve this with caches — keep data close. But we threw out the big caches when designing our GPU. So what do we do? The solution is beautiful. Instead of hiding the latency — the GPU **fills** it with useful work. On a single processing block (SM — Streaming Multiprocessor) the GPU keeps register state for not one warp, but **dozens**. Switching takes a single cycle, zero overhead. Warp A requested a texture and is waiting? The GPU switches to warp B. Then C. Then D. By the time data for A arrives from memory — its turn has come back around. To eliminate stalls, the GPU simply does other work while waiting. The more warps in flight — the better the latency is hidden. 200 cycles of latency, one warp computes for 60 — so you need at least 4 in the queue. In practice a GPU keeps 32–64 warps on a single SM and almost never idles. Key principle — remember this one. > CPUs hide latency with caches — keep data nearby. GPUs — by switching threads: while one waits, another computes. More on GPU architecture, warps, and latency hiding — in the free [Cornell Virtual Workshop: Understanding GPU Architecture](https://cvw.cac.cornell.edu/index) course. ### We just built a GPU Let's look back. In three steps we designed a new type of processor. Threw out everything the CPU had that's unnecessary for graphics — branch prediction, instruction reordering, huge caches. Packed the die area with dozens of simple ALUs. Grouped them with a shared decoder — SIMD. Solved the slow memory problem by switching between threads. That's a GPU. This is exactly how NVIDIA engineers reasoned in the late 90s. Put it all together — and in 1999 the GeForce 256 arrives. They themselves called it "the world's first GPU." 17 million transistors, 120 MHz, 10 million polygons per second. The key feature — hardware Transform & Lighting: vertex transformations and lighting calculations moved entirely from the CPU to the graphics card for the first time. One chip — the entire rendering pipeline. ![NVIDIA GeForce 256 — первый в мире GPU, 1999](https://anvilarth.github.io/geforce256.png) _NVIDIA GeForce 256 (1999) · 17M transistors · 120 MHz · 10M polygons/sec · photo: Wikimedia Commons, public domain_ ## From Shaders to CUDA ### Shaders: from hardwired to programmable We hit compute speed hard — more transistors, more ALUs, beautiful. But there was another problem: everything was hardwired into silicon. Want a different effect? Wait for a new chip. Why hardwire at all? Transistor budget — 17 million, every one counts. Here's a concrete example. Take a single pixel on a wall in Quake near a torch. To compute its color, the chip runs the pixel through a chain of hardwired operations: The problem is specific. That formula we just saw — it computes wall lighting. Light falls, reflects, you get a specular highlight. For a brick wall in Quake — perfect. But water reflects and refracts light completely differently. Skin scatters light below the surface. Metal gives a sharp highlight, fur gives a soft one. Each material needs its own formula. And the chip has just one hardwired in. And you can't hardwire all variants into silicon — there are way too many. But transistor count growth helped: 17M (1999) → 57M (2001) → 107M (2002). There was now budget for experimentation. And that's how **shaders** appeared — small programs written by the developer that the GPU executes for every vertex or pixel. The difference is fundamental: in fixed-function, transistors are wired together in a specific order — the chip always runs exactly the chain that was soldered at the factory. In the shader model, the chip has a universal ALU + instruction memory. Before rendering, you load your program — any program. The same ALU can compute walls, water, and skin — just load a different formula. They scaled up: _[схема] Growth in programmability: from 128 instructions with no branching to a full-fledged processor in 3 years._ **Shader Model 1.0** (2001, GeForce 3) — up to 128 instructions, no branching. Basically the same fixed-function, except you could rearrange the steps. The programmable block was 2-3x slower than the hardwired one — flexibility costs a lot. **Shader Model 2.0** (2002, Radeon 9700) — 256 instructions, `if/else` for the first time. You could now write real logic, less overhead. **Shader Model 3.0** (2004, GeForce 6) — up to 65,000 instructions, loops, dynamic branching. A full-fledged processor. Overhead got so small that drivers started emulating fixed-function via shaders — the hardwired blocks literally lost their purpose. Why the hell hardwire operations into silicon if software is almost as fast and can do a million more things? Flexibility won. > The GPU went from a "calculator with buttons" to a "programmable calculator". _[схема] Unified architecture: instead of 8V+24P -- 128 universal blocks with dynamic allocation._ The next logical step — all blocks on the chip became identical. 128 universal blocks, the scheduler distributes tasks on its own. Utilization close to 100%. And then NVIDIA engineers realized something even more interesting. If all 128 blocks can do any math — it's no longer a graphics card. It's a **massively parallel general-purpose compute engine**. Gaming optimization accidentally created a supercomputer on a graphics card. ### CUDA: how to even get at this thing Ok, we've got a supercomputer on a graphics card. But the only way in was through a graphics API — OpenGL or DirectX. Want to solve an equation? Encode your data as a texture, wrap your computation as a shader, read the result back as an image. A scientist literally pretends to draw a picture so the GPU computes physics for them. Reminds me of [Infinite Storage Glitch](https://github.com/KKarmugil/Infinite_Storage_Glitch) — a project where folks encode files into video pixels and upload them to YouTube as free cloud storage. A video platform, but technically streaming arbitrary bytes. The solution came pretty fast. In 2004, Ian Buck at Stanford built Brook — the first GPU language without a graphics API. Clunky, limited — but proved the idea works. NVIDIA hired him, and by 2007 **CUDA** shipped. What changed fundamentally? Instead of "encode your data as a texture, pretend you're drawing" — three simple ideas: **1. Just write a function** — mark it `__global__` and say "launch on 10 000 threads". That's it. The GPU distributes across cores on its own. **2. Thread hierarchy** — threads group into blocks, blocks into a grid. You think about task structure, not hardware. Remember warps from Chapter II? Here's how it connects: you say "a block of 256 threads", and the GPU internally slices it into warps of 32 — those same groups with a shared decoder and SIMD. The programmer manages blocks, the GPU manages warps. **3. Threads communicate** — inside a block there's fast shared memory. Threads can exchange data directly, without routing through slow global memory. In practice — you write normal C code, mark the function `__global__`, compile with `nvcc`, and run it. The same 128 cores crunch your task. No textures, no triangles, no pretending. Just math. The hardware didn't change — what changed is how you talk to it. And scientists caught on fast, because the pattern is the same as in graphics: one formula across millions of data points. The GPU became a universal compute engine. But every new workload exposed weak spots — and demanded new solutions in hardware. ## From compute engine to AI chip Ok, we have a universal parallel compute engine with a decent interface. Victory? Not quite. Every time scientists and engineers started actually using GPUs — they found a new problem. And every architecture generation is a response to a specific pain point of the previous one. ### Fermi → Maxwell: GPU learns to be reliable (2010–2014) The first problem was unexpected. Scientists ran a physics simulation — and the GPU result differed from the CPU. Not because the algorithm was wrong, but because a bit randomly flipped in memory. For games that's invisible — a pixel blinked, who cares. For science — a catastrophe, the calculation is wrong. ### Pascal: one card isn't enough (2016) Models grew faster than memory. One card — 12–16 GB, but you need more. What do you do? Put several side by side. But they talked over PCIe — a universal motherboard slot that connects literally everything. Data went through the CPU: GPU → CPU → GPU, just ~16 GB/s. Pascal made a direct GPU-GPU link without the detour through CPU — NVLink, 160 GB/s, 10x wider. Multiple cards started working almost as one. ### Volta: specialization comes back (2017) Now here's the legendary twist. Remember, in Chapter III we removed fixed-function operations from GPU for flexibility? Now NVIDIA brings them back — but for a different task. ML is essentially matrix multiplication, millions of times in a row. A regular CUDA core does one operation per clock. A whole 4×4 matrix — 64 clocks. Wasteful. ### Ampere: two worlds on one chip (2020) Tensor Cores worked great, but the world changed — models with 175 billion parameters appeared (GPT-3). And it turned out that training and inference are completely different workloads. Training wants maximum compute, inference wants minimum latency. One chip has to handle both. Ampere A100 solved this with three tricks. **TF32** — computes fast like FP16, but nearly as precise as FP32, and crucially without code changes. **Sparsity 2:4** — if half the matrix is zeros, GPU skips them in hardware, 2× speedup for free. **MIG** — one GPU can be sliced into 7 isolated chunks, each running its own workload. The first chip designed for both training and inference. ### Hopper: the era of large models (2022) LLMs — GPT, LLaMA — turned out to be hungry in a whole new way. Different model layers need different precision, but FP16 everywhere is wasteful: we're spending memory and bandwidth on precision that isn't needed. ### Blackwell: a chip can't grow forever (2024) A single monolithic chip hit the physics wall — the bigger the die, the more defects during manufacturing. Plus 700W per chip — a data center with thousands of these consumes as much power as a small town. ### What's next: Feynman (2028) NVIDIA already showed the roadmap. The next big step is Feynman (2028). No room to grow horizontally — so we grow vertically. ### The entire evolution at a glance ## From drawing pixels to almost AGI We've traveled from a thing that draws triangles in Quake to chips that train models with hundreds of billions of parameters. The GPU started as a specialized accelerator for a single task, became a general-purpose compute engine — and now it's specializing again, only this time for matrix multiplications. And it's precisely this new specialization that raised the question: maybe the GPU isn't the only answer? ### The inference problem When an LLM generates text — one token at a time. For each token the model reads all its weights from memory: billions of parameters. Compute per byte read — tiny. To fully utilize an H100 you need ~300 ops per byte read, but in practice you get ~1. Thousands of cores and Tensor Cores sit idle, waiting for data from HBM. Utilization 5–15%. You bought a $30K card and you're using 10% of it — that hurts. > The core problem: **how to stop waiting on memory**. TPU, Groq, Cerebras — they all solve exactly this, but in different ways. A GPU has thousands of cores, and each one goes to shared memory for data on every step. Computed — wrote back. Computed — wrote back. Universal, but the cores are constantly waiting for memory to respond. ### Google TPU: data flows instead of sitting still The TPU is fundamentally different. Let's walk through it step by step. In a GPU every core goes to memory on every step: fetched a number, multiplied, wrote back, fetched the next one. In a TPU it works like this: **Step 1.** We load weights from memory into the array — once. Each element in the array stores its own weight. Element [0,0] stored W₁, element [0,1] stored W₂, and so on. The weights are "frozen" in the array. **Step 2.** We feed input data X₁ from the left into the first element. It multiplies X₁ by its W₁ and passes the result to its neighbor on the right. Not to memory — directly to its neighbor. **Step 3.** The neighbor receives the result, adds its own W₂×X₂, passes it along. The next one — same thing. On the right side out comes the finished result of a matrix row. **Step 4.** While the first row is "passing" through the array, the second one is already being fed in behind it. The elements don't idle for a single cycle. Bottom line: GPU hits memory **on every operation**. TPU — **once for the entire matrix**. The difference in memory accesses is massive, and that's exactly why TPU is faster for matrix multiplications. Google's main bet is scale: TPU Pods with thousands of chips connected by a fast network. The upside: if you're in the Google Cloud ecosystem — it can be cheaper and faster than GPUs. But you're locked into Google, no running your own. ### Groq LPU: removed memory altogether Groq solved the problem even more radically. Remember from Chapter II — HBM is external memory, sits next to the chip but is physically separate. Every access — hundreds of cycles of latency. SRAM — memory right inside the chip, on the same die, accessible in 1-2 cycles. Groq figured: if during inference we just read weights sequentially and the model fits in SRAM — why go to external memory at all? They removed HBM, everything's on-chip. Latency vanished, inference speed became deterministic. But the limitation is obvious: SRAM is finite, large models don't fit. ### Cerebras WSE: a chip the size of a wafer Cerebras came at it from a third angle: if the problem is that data is far from compute — make the chip gigantic so everything fits. One die the size of an entire silicon wafer — 850,000 cores, 44 GB SRAM. Why slice a wafer into small chips and then connect them back together if you can just leave it whole? The upside: incredible density. But manufacturing is hard and the ecosystem is small. ### What this all means At the start of this post the GPU was a black box to me. Throw in a model, press a button, wait. Not anymore. We've gone from drawing triangles in Quake to chips that train models bringing us closer to AGI. And behind every solution — not magic, but a specific engineering problem and a specific tradeoff. My main takeaway is this: when you break something down to first principles — you stop being afraid of complexity. The GPU seemed impossibly tangled, but turned out to be a set of clever solutions to physics constraints. Parallelism — because pixels are independent. SIMD — because the formula is the same. Latency hiding — because memory is slow. Tensor Cores — because ML is matrices. Every technology works this way — break it down to first principles and it clicks. And that's a skill that transfers to everything else. But knowing the architecture is half the battle. The other half — writing code that actually takes advantage of it. How to write a CUDA kernel yourself. What libraries and tools are out there. How to tell that your code is using 10% of the GPU, and what to do about it. That's what the next post is about. ### Sources - [CMU 15-462 — How a GPU Works](https://www.cs.cmu.edu/afs/cs/academic/class/15462-f11/www/lec_slides/lec19.pdf) - [Fabian Giesen — A trip through the Graphics Pipeline](https://fgiesen.wordpress.com/2011/07/09/a-trip-through-the-graphics-pipeline-2011-index/) - [CMU 15-418 — GPU Architecture](http://15418.courses.cs.cmu.edu/spring2015/lecture/gpuarch/slide_003) - [NVIDIA Blog — 25th Anniversary of GeForce 256](https://blogs.nvidia.com/blog/first-gpu-gaming-ai/) - [Wikipedia — CUDA](https://en.wikipedia.org/wiki/CUDA) - [Yahoo Finance — Going all-in with Nvidia](https://finance.yahoo.com/news/going-all-in-with-nvidia-how-jensen-huangs-high-stakes-bets-paid-off-113053891.html) - [The Chip Letter — Nvidia's Embarrassingly Parallel Success](https://thechipletter.substack.com/p/nvidias-embarrassingly-parallel-success) - [Acquired Podcast — Nvidia Part I](https://www.acquired.fm/episodes/nvidia-the-gpu-company-1993-2006) - [CloudFleet — NVIDIA GPU Architectures](https://cloudfleet.ai/blog/cloud-native-how-to/2023-03-comparison-of-different-nvidia-gpu-rchitectures/) - [Google — In-Datacenter Performance Analysis of a Tensor Processing Unit (TPU paper)](https://arxiv.org/abs/1704.04760) - [Pope et al. — Efficiently Scaling Transformer Inference](https://arxiv.org/abs/2211.05102) - [NVIDIA — Volta Architecture Whitepaper (Tensor Cores)](https://images.nvidia.com/content/volta-architecture/pdf/volta-architecture-whitepaper.pdf) - [NVIDIA — Hopper Architecture Whitepaper (H100, Transformer Engine)](https://resources.nvidia.com/en-us-grace-hopper/gtc24-whitepaper-hopper) - [NVIDIA — Blackwell Architecture Whitepaper (B200, 208B transistors)](https://resources.nvidia.com/en-us-blackwell-architecture) - [NVIDIA — NVLink & NVSwitch](https://www.nvidia.com/en-us/data-center/nvlink/) - [Wikipedia — Brook (Ian Buck, Stanford GPGPU)](https://en.wikipedia.org/wiki/Brook_(programming_language)) - [Wikipedia — Shader History (Shader Model evolution)](https://en.wikipedia.org/wiki/Shader#History) - [Groq — LPU Architecture (SRAM-only, deterministic inference)](https://groq.com/technology/) - [Cerebras — Wafer-Scale Engine (850K cores, 44GB SRAM)](https://www.cerebras.net/chip/) - [Wikipedia — Systolic Array](https://en.wikipedia.org/wiki/Systolic_array) --- # Age of Autoresearch _Andrey Filatov · 2026 · ~15 min read_ _Source: https://anvilarth.github.io/autoresearch.html · Author: Andrei Filatov (https://anvilarth.github.io/)_ ## Chapter 1. A Lazy Evening It was an ordinary weekday evening, and I was honestly too lazy to test hypotheses by hand. Not lazy as in "tired and want to sleep" — lazy in a specific, engineering sense: I had a backlog of fifteen-ish ideas that needed to be run, checked, filtered for garbage, repeated. Work I'd done dozens of times and knew every step of by heart. And that evening I just didn't have it in me to sit down and click through the same thing manually one more time. For the past month and a half I'd been living without weekends, trying to become a sort of 10x engineer — doing a lot more without raising my own cognitive load. Sounds nice. In practice it meant I was trying to shove agents into everything and hoping it would somehow speed itself up. The first attempt was brazen and dumb at the same time: I just launched ten different tasks in parallel, agent upon agent. The logic was simple — if one agent handles one task fine, ten agents will handle ten tasks at once, and I get my evening back. In practice, almost always one of the ten broke. And the moment it broke, all of my attention got stuck on it. Net result: I spent more time cleaning up than if I'd just done them one by one myself. The aftertaste was mega unpleasant: I thought I'd split the work into parallel streams and instead got parallel chaos. The second attempt was smarter. I figured — okay, if an agent without an explicit plan produces nonsense, I'll give it a plan. I wrote the agenda up front: what to do in which order, which corner cases to watch for, what counts as success. The agent behaved noticeably cleaner — fewer surprises, more cases covered. Things got better. But still kind of lousy — I was still sitting next to it cleaning up the loose ends, there were just fewer of them. It felt like I had improved the supervision process rather than gotten rid of the need to supervise. And so on that lazy evening, instead of trying a third time to steer the agent better by hand, I did a lazy thing. I simply described my hypothesis-generation pipeline in text — how I usually come up with what to try next — and asked Gemini to evaluate the result of each attempt. Not to steer the process, not to decide the next step, but specifically to judge: here's a hypothesis, here's the result, is it good enough or not. I handed the whole thing to the agent and went to bed, because I was honestly too lazy to sit there and watch. In the morning I opened the log and found that **40 hypotheses** had been tested overnight. ![Agent terminal: Pursuing goal (1d 1h 41m)](https://anvilarth.github.io/img/sankalp_17pm.webp) _This is what it looks like from the outside: a line in the terminal and a counter ticking while you sleep. The screenshot isn't mine — it's from [sankalp's write-up](https://sankalp.bearblog.dev/autoresearch/), which is what the next chapter is about._ Not all forty turned out useful — some were outright garbage, some repeated what I'd already tried. But that didn't matter. All night I did nothing by hand — and came out with more tested variants than I'd have made in a week with a laptop. And it became clear: the point wasn't to steer the agent better. The point was who evaluates the result. It worked once, at night, on one specific task. What remained was to understand what exactly had worked — and oddly enough, what helps you understand that isn't your own experience, but the fact that someone else had stumbled onto almost the same thing a bit before me. ## Chapter 2. It Already Had a Name What happened that night wasn't an invention. It already had a name: autoresearch. And a couple of loud precedents before me. First I came across [sankalp's post about QR decomposition](https://sankalp.bearblog.dev/autoresearch/). A classic task: decompose a matrix fast. He didn't touch the kernels by hand — he set Codex loose on the problem as an autonomous agent. It kept a beam of candidates, spawned sub-agents for profiling, math, and idea generation, and culled the dumb branches on its own as it went. The result — a **232x speedup** on batched QR and 12th place out of 183 participants, despite the author having no professional GPU experience. ![Beam search diagram: candidate branches, merge points, best branch](https://anvilarth.github.io/img/sankalp_10pm.webp) _sankalp's beam of candidates: branches live in parallel, weak ones fade, strong ones merge. Nobody picks the winner by hand — the criterion picks it. Diagram from [sankalp's post](https://sankalp.bearblog.dev/autoresearch/)._ And as I read, it clicked that it wasn't about the agent or the prompt. Before, the expert invested in the solution itself: assembling features by hand, inventing heuristics from their head, writing kernels natively. Knowing how the problem worked was the job. Now that knowledge is almost unnecessary directly. What you need is the ability to build a cycle that finds the solution itself: benchmark, oracle, stopping criterion. You don't write the kernel — you build the loop that searches through it. The same _bitter lesson_, just not about model architecture, but about what the human is busy with. Two direct conclusions follow. First: you can physically do less work. That very hypothesis-testing routine I was too lazy to do in the evening gets delegated — the agent runs it at night, I sleep. Second: the work gets done better. Not "the same thing, faster," but qualitatively better — in the same time, the oracle searches the solution space wider than a single person and finds variants I simply would never have reached on my own. And sankalp isn't the only one. Karpathy put together [autoresearch](https://github.com/karpathy/autoresearch): a 630-line script, an autonomous loop — the agent reads the training code, proposes a change, runs a five-minute run, measures the improvement, commits or rolls back, and goes again. Overnight — hundreds of experiments, not a single click. [AlphaEvolve](https://deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/) from DeepMind — the same idea at a larger scale: Gemini plus an evolutionary framework plus automated evaluators. They improved solutions for 20% of 50 open mathematical problems, and sped up datacenters, chip design, and AI training at Google. ![autoresearch progress chart: 83 experiments, 15 kept improvements](https://anvilarth.github.io/img/karpathy_progress.png) _One night in Karpathy's repo: 83 experiments, 15 of them kept improvements. The gray dots are what the loop culled on its own. Exactly the garbage proportion I saw in my log that morning. Chart from [karpathy/autoresearch](https://github.com/karpathy/autoresearch)._ In Karpathy's README there's a line that describes this shift more precisely than I'd formulated it myself: > You're not touching any of the Python files like you normally would as a researcher. Instead, you are programming the `program.md` Markdown files that provide context to the AI agents and set up your autonomous research org. The file the agent edits and the file the human edits are different files. All of the researcher's work has moved into the second one. Different names, different implementations, different scale. But the same pattern: expertise moves from "find the solution" to "build the system that finds the solution." And so far that's a frame, not an answer. To understand where the frame's edges are, what helped me wasn't other people's cases but two of my own. One went smoothly — there's almost nothing to tell about it. The second — the one where I had to argue with myself: I was sure I knew better. ## Chapter 3. I Knew Better For two weeks I had been building dataset filters by hand. A simple-looking task — decide which samples to throw out of training and which to keep. In practice, a mix of engineering and curation: you read papers, look at the data, formulate a hypothesis, test it, throw it away, repeat. I believed I understood this task. Armed with ego and knowledge of how to do it better, I spent two weeks pouring intuition and paper-reading into manual filters. And then I did exactly what I did on that lazy evening — built a benchmark with Gemini as the oracle and set a zero-shot VLM loose on it. Just to see whether it would beat my manual work. No heuristics or hints, no involvement from me. The first attempt was already better than two weeks of manual work. The second — even better. It wasn't some cosmic gap. Not "the agent invented a filter I couldn't have thought of." The result was simply objectively better, and the time it took — a couple of runs versus two weeks of my life. And that's when it got genuinely painful — not because the method worked, but because I'd spent two weeks arguing with myself over a solution that the automation searched through in a couple of attempts. Before AlexNet, features were assembled by hand. Engineers invented heuristics, picked out features, tuned pipelines — and that was expertise backed by years of work. Neural networks automated it: features became learned, and the work shifted to architecture and data. Now the researcher's expertise itself is being automated. You don't need to know the domain perfectly to make good decisions — it's enough to build an oracle that knows what counts as a good solution and let the agent run. That's the bitter lesson on my own skin. Two weeks beaten by a single night. The second case went down much more calmly. Maybe the stakes were different — or maybe after the first one I stopped arguing with the oracle. ## Chapter 4. Not a Silver Bullet The second case was about speeding up training. No inner drama — just a task where frontier models are expensive and can be overkill, but if you point them at optimizing inference or training time, they save money and time. I left the model working for two days and got a **30% speedup**. Pushing further got much harder, and there was no point spending more time on it. Smooth, mundane, no ego story. But even on a smooth case — and especially afterward — pitfalls surfaced, the kind that mean this method shouldn't be treated as a silver bullet. The first one — loss parity. I forgot to add a check to the oracle that the loss doesn't drift apart. The agent honestly reported "done, there's a speedup" — when in fact the loss was NaN everywhere. The model had formally become faster, but it had stopped learning. Once I added the loss parity check, the problem went away. A banal thing that I nonetheless missed, because I trusted the agent too much to decide what counts as success. The second — an insufficiently detailed oracle. This surfaced later, on a text-removal-from-images task. I was testing different models but validating the result on the full image. Small artifacts — leftover letters, a half-erased character — I noticed with my eyes, but the oracle didn't. It considered the text removed, and the solution it called best in fact left garbage behind. As soon as I added a crop of the region where the text was supposed to be removed to the oracle — validating not the whole picture but the actual fact of removal — a completely different pipeline turned out to be the best. The previous winner had simply been exploiting the oracle's blindness. ![Chart of one session on KernelBench-Mega: speedup relative to reference versus tokens spent](https://anvilarth.github.io/img/kernelbench_session.jpg) _What a properly built session looks like — a breakdown of one run on [KernelBench-Mega](https://kernelbench.com/mega). For the first 64% of the time there's no code at all: measuring the baseline, microbenchmarking barriers, deriving the roofline. The first working kernel appears at the 224k token mark. And near the end — the point I'm showing this for: "finer split-K regresses → measured, reverted". The hypothesis didn't work, the measurement caught it, rollback._ The benchmark's author describes that same run in one sentence: > It spent 64% of the session in silence timing the baseline, microbenchmarking grid barriers, deriving a ~29x bytes/token roofline. The one regression it tried (finer split-K) it measured and reverted instead of rationalizing. Measured and reverted instead of rationalizing — that's a working oracle. My NaN happened precisely because there was nothing to measure with, and rationalizing was all the agent had left. The general lesson is simple and unpleasant: engineering autoresearch is a delicate task, where time should be spent specifically on designing the oracle and the loop. The time doesn't go into steering the agent or hunting for a magic prompt — it goes into honestly and thoroughly describing what counts as a good solution. Treating this as an automatic silver bullet means sooner or later getting a NaN where your loss used to be. And even when the oracle is built right and the loop runs smoothly, one variable remains that design can't remove — the model inside the loop itself. I ran this whole finely-tuned process on different models, and it turned out that with the same oracle and the same task, the result depended not only on what I'd built, but on whom I'd entrusted it to. ## Chapter 5. Not Power, but Intuition On the code-speedup task I had three models: Fable 5, Claude Opus 4.8, and GPT-5.5. Same task, same oracle. Fable 5 — I just described the task, and everything worked. The agent produced hypotheses that made sense, tested them, culled the dumb ones. All I had left to do was nod at the reasonable steps. Claude Opus 4.8 proposed a solution that caused an OOM, and sort of gave up at that point — didn't correct itself, didn't try another path. GPT-5.5 eventually managed, but I had to explicitly point it where to look, otherwise it just treaded water. The difference isn't in the ability to execute steps. All three can read code, write code, run, measure. The difference is in domain intuition — in how deeply the understanding of what's even worth trying in this task is baked into the model. Fable 5 had that understanding. The other two didn't. I compensate for models' lack of intuition by throwing context into the prompt — blogs, guides, other people's cases. But that's a limitation, and an unpleasant one: to know what to help the model with, I have to already understand the topic well myself. In other words, the very expertise I thought I was automating, I have to keep in my head so I can pull it out in pieces and feed it to the agent. And this isn't only my observation. The same pattern shows up in independent benchmarks. [KernelBench-Mega](https://kernelbench.com/mega) is a benchmark for whole-block megakernels, where you need to fuse an entire model block into one kernel rather than optimize individual operations. All previous models "won" on it only through a multi-kernel Triton pipeline — technically passes the test, but there's no honest fusion, it's a workaround. And only Fable 5, [according to the benchmark author's tweet](https://x.com/elliotarledge/status/2072814573753975266), wrote the first real megakernel. A direct parallel with my pitfall from chapter four: a poorly designed oracle allowed sleight of hand until someone did it honestly. [AutoKaggle](https://arxiv.org/abs/2410.20424) — the same pattern outside GPU: a multi-agent framework for Kaggle competitions, where the oracle is the competition itself and the tests, and across eight competitions it reached a 0.85 validation submission rate and 0.82 comprehensive score, at human level. ![Elliot Arledge's tweet about the first real megakernel](https://anvilarth.github.io/img/tweet_arledge.png) _The benchmark author lists who "won" before and by how much: Opus 4.8 — 14.4x, GLM-5.2 — 11.1x, GPT-5.5 — 4.3x, Sonnet 5 — 4.0x. All through a multi-kernel Triton pipeline that doesn't pass the authenticity gate. [Full tweet](https://x.com/elliotarledge/status/2072814573753975266)._ So the best tool isn't the most powerful model — it's the model with the right intuition for the specific task. And the more precisely the oracle is built, the clearer you can see who has that intuition and who has to compensate for it with context. Where this is heading is recursive loops — cycles where the agent improves not the model's weights but the system itself: the code, the pipeline, the artifact. Each improvement makes the next one easier. I hope tools you can simply use will appear soon. For now you have to pick which loop fits which task every time, and it works only in spots. So what does that say about my own role, if intuition — the only thing that's still mine — is becoming a commodity you can just throw into a prompt? ## Chapter 6. What's Left To the question "will AI replace us" I now have a boring answer: for part of the work, it already happened. Interns and juniors used to draw charts and test simple hypotheses — now the agent does all of that. But that's about juniors. With me it's more complicated. Three years ago I spent three days digging through PyTorch Lightning's source code to understand how checkpointing works there. Three days — and I knew it at a level where I could explain it to anyone and fix anything. Now an agent does the same in an hour. Frees you from the routine — yes. But takes something in return. That feeling of authorship that hand-digging through source code used to give — "I understood this, I did this" — it's gone. Factually, I did nothing, the agents did it all. It's like moving from a senior role to a manager role. The story is exactly the same as in an ordinary career: you become a team lead — and you barely do anything with your own hands anymore. Before, you sat and solved things yourself; now you watch others solve them and make sure they solve them in the right direction. The output is more than you'd have made alone. But there's nothing hand-made by you in it anymore. The position is less dopaminergic: the old high from solving something with your own hands won't be there. And I thought I could stop there. A manager is still needed, after all. Not because he's smarter than the people he manages, but because he knows where to point them — he has domain intuition. That was my conclusion in the previous chapter: the difference isn't power, it's intuition. Where the model lacks intuition, I compensate by throwing context into the prompt. Intuition was the last thing that was still mine. And then, over a couple of months, several stories piled up that don't fit this picture. The Jacobian Conjecture — a classic conjecture in algebra. If a polynomial map has a nonzero constant Jacobian, then the map is invertible. Recently — a counterexample in dimension three. And it wasn't found by a group of mathematicians. It was found by Fable 5, while I was watching the World Cup final. Terence Tao, a Fields Medal laureate, wrote a [post](https://terrytao.wordpress.com/2026/07/21/a-digestion-of-the-jacobian-conjecture-counterexample/) to digest this. He explains retrospectively, geometrically, what the model found on its own. And here's what struck me: Tao sits down and counts just how improbable this is. A degree-seven polynomial. Its Jacobian has a priori up to **1329** nonzero coefficients — against **360** degrees of freedom of a general polynomial map of the same degree. 1329 equations that must vanish simultaneously, in a space of 360 variables. You don't find something like that by brute force — the chance of hitting the right point is zero. So Fable wasn't searching. It saw the structure. ![A paragraph from Terence Tao's post counting 1329 equations against 360 degrees of freedom](https://anvilarth.github.io/img/tao_1329.png) _The very paragraph from [Tao](https://terrytao.wordpress.com/2026/07/21/a-digestion-of-the-jacobian-conjecture-counterexample/). The last sentence is a death sentence for brute force: "finding such a polynomial looks highly unlikely to be located by brute force"._ At the end of the post there's a disclaimer that five years ago would have looked wild in the text of a Fields medalist: > I used an AI chatbot to discuss various aspects of this problem and to confirm several of the calculations made here. Seeing where the solution lies in a space where naive search is useless. That's exactly what I considered mine. In university they told me there are different levels of knowing the material. The first stage — you learn the terms: "Jacobian", "polynomial", "invertible map". You know they exist, but you don't understand them. The second — you know all the material: definitions, theorems, proofs. But you can't apply it — you bomb the exam problem because it differs slightly from the template. And there's a third stage, which is what actually knowing the material means — intuition. Not just knowing that a fact exists, but seeing when to apply it, and recognizing the same structure in a new situation. Models have walked the same path. First stochastic parrots — guessing the next word. Then they learned the material, know the facts, but break on anything slightly new. And now — the third stage. They see structure, find solutions that weren't in the training data. A counterexample to a conjecture isn't a reproduction of something the model saw, it's new mathematics. And there's no cheating here: a counterexample either works or it doesn't — in pure mathematics there's no such thing as a badly designed oracle you can route around. Fine, let's say the model has intuition. But the manager who points the way is still needed — that's what I was holding on to. The second case is exactly about this, and it's closer to me, because you can see what the human did there. Mathematician Dmitry Rybin took the Dinits–Garg–Goemans conjecture — a problem about network flows, open since the nineties: can a fractional solution be made unsplittable without raising the cost and without overloading any edge by more than the size of the largest shipment. The counterexample was found by [GPT-5.6 Pro](https://chatgpt.com/share/6a60b2eb-0b64-83ee-9c76-7931ca1de063): a graph on seven nodes, three shipments, fractional cost 58, while any unsplittable routing costs at least 60. ![Counterexample to the Dinits–Garg–Goemans conjecture: a graph on seven nodes with flows, costs, and demands](https://anvilarth.github.io/img/rybin_graph.jpg) _The entire counterexample in full: seven nodes, one source s, three sinks with demands 15, 10, and 15. Blue numbers are flow, red are edge costs. For thirty years nobody could either construct such a graph or prove it doesn't exist. [Image from Rybin's tweet](https://x.com/DmitryRybin1/status/2079904005652893709)._ What's interesting here isn't the model but the prompts. There were four of them, under sixty words in total: "construct a counterexample", "you should do a breakthrough", "it's enough of partial results, let's finish". Three sessions in a row, an hour and a half each, the model returned partial results, and the human simply asked it to continue. There is zero domain intuition in those words. There's direction, and an understanding of when the result is already good enough — that is, exactly what I'd kept for myself as the manager. Turns out, that's about a minute's worth of work. The same skill transfers to my job too. My personal labor — what is it? A set of patterns I've built up over the years. "I've seen this bug — I know where to look." "This architecture breaks like this — route around it like this." "This hypothesis won't work because this metric lies." Those are all patterns. And patterns can be learned. And a network that found a solution in a space of 1329 equations — it will learn my patterns too. The only question is the price, and that price is cheaper every month than it was yesterday. And here's the strange thing: the joy hasn't gone anywhere. I still open the log in the morning and stare at what happened overnight, even though I did nothing by hand. So the joy was never about "I solved it myself" — it was about the problem moving. And it moves a lot more now: I pull most of the model training alone, where a team used to be needed. So I'm not quitting — I'm building the next loop. But to the question of where my value is now, I have no answer. ### Sources - [sankalp — Auto-Research: QR Decomposition](https://sankalp.bearblog.dev/autoresearch/) — 232x speedup on batched QR (419,000 µs → 1805 µs) via Codex as an autonomous agent; 12th place among 183 participants, more than 1500 submissions in 14 days. - [Andrej Karpathy — autoresearch](https://github.com/karpathy/autoresearch) — a 630-line script: training code → change → 5-minute run → improvement measure → commit/rollback → repeat. - [DeepMind — AlphaEvolve](https://deepmind.google/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/) — Gemini + evolutionary framework + automated evaluators; improved solutions for 20% of 50 open mathematical problems. - [Terence Tao — A digestion of the Jacobian conjecture counterexample](https://terrytao.wordpress.com/2026/07/21/a-digestion-of-the-jacobian-conjecture-counterexample/) — counterexample in dimension 3; degree-7 polynomial, up to 1329 Jacobian coefficients against 360 degrees of freedom. - [KernelBench-Mega](https://kernelbench.com/mega) — independent benchmark for whole-block megakernels (author: Elliot Arledge). - [Elliot Arledge (tweet)](https://x.com/elliotarledge/status/2072814573753975266) — the first real megakernel on KernelBench-Mega, written by Fable 5. - [AutoKaggle (arXiv:2410.20424)](https://arxiv.org/abs/2410.20424) — multi-agent framework for Kaggle competitions; 0.85 validation submission rate, 0.82 comprehensive score across 8 competitions. - [Dmitry Rybin — counterexample to the Dinits–Garg–Goemans conjecture](https://chatgpt.com/share/6a60b2eb-0b64-83ee-9c76-7931ca1de063) — four prompts, under 60 words; three sessions with partial results, counterexample on the fourth: a graph on 7 nodes, three shipments, fractional cost 58 versus a minimum of 60 for any unsplittable routing.