Телеграм в третий раз за 2 недели стёр пост из черновиков, поэтому поста не будет 🤷♂️ Ещё раз — и пишу Дурову 👶 Держите ссылку https://epoch.ai/gradient-updates/ai-progress-is-about-to-speed-up и тезисы на англисйком: > The release of GPT-4 in March 2023 stands out because GPT-4 represented a 10x compute scale-up over the models we had seen before. Since then, we’ve not seen another scale-up of this magnitude: all currently available frontier models, with the exception of Grok 3, have been trained on a compute budget similar to GPT-4 or less > Grok 3 represent more than an order of magnitude scale-up over GPT-4, and perhaps two orders of magnitude when it comes to reasoning RL. Based on past experience with scaling, we should expect this to lead to a significant jump in performance, at least as big as the jump from GPT-3.5 to GPT-4. > The models are initially going to be perhaps an order of magnitude bigger than GPT-4o in total parameter count so we’ll probably see a 2-3x increase in the API token prices and around 2x slowdown in short context decoding speed when the models are first released, though these will improve later in the year thanks to inference clusters switching to newer hardware and continuing algorithmic progress. > What should we make of Grok 3? It’s possible to make both a bullish and a bearish for scaling based on Grok 3. The bullish case is that Grok 3 is indeed state-of-the-art as a base model with a meaningful margin between it and the second best models, and this is what we would expect given its status as a “next generation model” with around 3e26 FLOP of training compute. The bearish case is that the gap between Grok 3 and models such as Claude 3.5 Sonnet seems much smaller than the gap between GPT-4 and GPT-3.5, despite both representing roughly an order of magnitude of compute difference. > I think the correct interpretation is that xAI is behind in algorithmic efficiency compared to labs such as OpenAI and Anthropic, and possibly even DeepSeek. This is why Grok 2 was not a frontier model despite using a comparable amount of compute to GPT-4, and this is also why Grok 3 is only “somewhat better” than the best frontier models despite using an order of magnitude more training compute than them. > Putting all of this together, I think Grok 3 gives us more reasons to be bullish than bearish on AI progress this year. === > In addition, a counterintuitive prediction I’m willing to make is that most of the economic value of AI systems, in 2025 and beyond, is actually going to come from these more mundane tasks that currently don’t get much attention in benchmarking and evaluations [про это я писал в канале в посте с критикой Gemini 2.0 Pro; ожидаю, что OpenAI смогут донести ценность]. The smaller improvements in long-context performance, ability to develop plans and adapt them to changing circumstances, a general ability to learn quickly from in-context mistakes and fix them, etc. are going to drive more revenue growth than the math, programming, question answering etc. capabilities that AI labs like to evaluate and demo.
Телеграм в третий раз за 2 недели стёр пост из черновиков, поэтому поста не…
Из этого канала
- #2349Чуть больше 2 лет назад узнал тут, что в США есть список запрещённых букв, с…
Чуть больше 2 лет назад узнал тут, что в США есть список запрещённых букв, с которых не может начинаться трёхбуквенное название аэропорта. Одна из них — Q.
- #2351(аххаха новую модель 3.7 назвали) anthropic.claude-3-7-sonnet-20250219-v1:0…
(аххаха новую модель 3.7 назвали) anthropic.claude-3-7-sonnet-20250219-v1:0 Claude 3.7 Sonnet is Anthropic's most intelligent model to date and the first…
- #2352Уже на claude.ai (даже для бесплатных пользвателей!) офф пост:…
Уже на claude.ai (даже для бесплатных пользвателей!) офф пост: https://www.anthropic.com/news/claude-3-7-sonnet
- #2347У этой работы есть ограничения, некоторые из которых плавно перетекают в намёки…
У этой работы есть ограничения, некоторые из которых плавно перетекают в намёки на то, что именно ждать от второй версии системы.
- #2344Картинки 1) устройство системы и описание того, как общаются агенты между собой…
Картинки 1) устройство системы и описание того, как общаются агенты между собой 2) Рост эло-рейтинга от количества времени работы системы (чем дольше работает,…