What Happens When AI Performance Asymptotes?
Tomasz TunguzVenture Capitalist If you were forwarded this newsletter, and you'd like to receive it in the future, subscribe here. What Happens When AI Performance Asymptotes?
In the past, the bigger the AI model, the better the performance. Across OpenAI’s models for example, parameters have grown by 1000x+ & performance has nearly tripled.
But model performance will soon asymptote - at least on this metric. This is a chart of many recent AI models’ performance according to a broadly accepted benchmark called MMLU. 1 MMLU measures the performance of an AI model compared to a high school student. I’ve categorized the models this way :
Over time, the performance is converging rapidly both across model sizes & across the model vendors. What happens when Facebook’s open-source model & Google’s closed-source model that powers Google.com & OpenAI’s models that power ChatGPT all work equally well? Computer scientists have been challenged distinguishing the relative performance of these models with many different tests. Users will be hard-pressed to do better. At that point, the value in the model layer should collapse. If a freely available open-source model is just as good as a paid one, why not use the free one? And if a smaller, less expensive to operate open-source model is nearly as good, why not use that one? The rapid growth of AI has fueled a surge of interest in the models themselves. But pretty quickly, the infrastructure layer should commoditize, just as it did in the cloud where three vendors command 65% market share : Amazon Web Services, Azure, & Google Cloud Platform. The applications & the developer tooling around the massive AI commodity brokers is the next phase of development - where product differentiation & distribution differentiate rather than brilliant, raw technical advances.2 1 MMLU measures 57 different tasks including math, history, computer science & other topics. It’s one measure of many & it’s not perfect - like any benchmark. There are others including the Elo system. Here’s an overview of the differences.. Each benchmark grades the model on a different spectrum : bias, mathematical reasoning are two other examples. |
Older messages
What If LLMs Change the Business Model of the Internet?
Wednesday, February 28, 2024
Tomasz Tunguz Venture Capitalist If you were forwarded this newsletter, and you'd like to receive it in the future, subscribe here. What If LLMs Change the Business Model of the Internet? Last
The Secrets to Building Vibrant Communities in Web3 Open-Source
Tuesday, February 27, 2024
Tomasz Tunguz Venture Capitalist If you were forwarded this newsletter, and you'd like to receive it in the future, subscribe here. The Secrets to Building Vibrant Communities in Web3 Open-Source
The First $100m ARR AI Security Company
Thursday, February 22, 2024
Tomasz Tunguz Venture Capitalist If you were forwarded this newsletter, and you'd like to receive it in the future, subscribe here. The First $100m ARR AI Security Company Palo Alto Networks,
Nobody Knows : Steel & Blockchains
Tuesday, February 20, 2024
Tomasz Tunguz Venture Capitalist If you were forwarded this newsletter, and you'd like to receive it in the future, subscribe here. Nobody Knows : Steel & Blockchains Asking “What problems
Building a $20b Behemoth : Office Hours with Steven Goldfeder of Offchain Labs
Monday, February 19, 2024
Tomasz Tunguz Venture Capitalist If you were forwarded this newsletter, and you'd like to receive it in the future, subscribe here. Building a $20b Behemoth : Office Hours with Steven Goldfeder
You Might Also Like
🚀 Lockheed Drops Terran Bid
Friday, May 3, 2024
Plus SES acquires Intelsat, Astroscale eyes IPO, $SPIR secures multi-million dollar deal and more! The latest space investing news and updates. View this email in your browser The Space Scoop Week
10words: Top picks from this week
Friday, May 3, 2024
Today's projects: JustSERP • MyArchitectAI • Fix My Speakers • Quizgecko • Airparser • ThoughtLeaders • GPT for Workspace • i Ask AI • Raguie • SaaS Dives • Language Atlas • AI Collection 10words
Descrb
Friday, May 3, 2024
Take photo, get product description, and many more BetaList BetaList Daily Descrb Take photo, get product description, and many more Too much email in your life? Switch to weekly emails or stop
🚀 Ready player one
Friday, May 3, 2024
Prepare to leave your 9-5 Dear , Ever watched "Ready Player One"? In that world, success in the OASIS hinges on knowing the game better than anyone else, mastering the clues, and leveraging a
Podscan, 2 months and $1k MRR in — The Bootstrapped Founder 317
Friday, May 3, 2024
2 months in. $1k MRR. And quite a few learnings.
Learity, Multilingue, and OneLLM.co
Friday, May 3, 2024
Create dashboards from your database with no SQL BetaList BetaList Daily Learity Exclusive Perk Create dashboards from your database with no SQL Multilingue Learn more than 20 languages the easiest way
We tested 7 AI image generators – with results
Friday, May 3, 2024
Plus tips, news & Buffer updates for your social media journey Image Hey there 👋🏾 If you've ever struggled with consistency in your content creation journey, you're not alone. I'
Paris’ je ne sais quant
Friday, May 3, 2024
Alice & Bob's quantum labs, Dubai's European founders & Goldman Sachs funds payments fintech View in browser Sponsor Card - Flagship-16 Good morning there, Many startups build their
How Texas’ online porn law could shatter a First Amendment precedent
Friday, May 3, 2024
After the Supreme Court declined to issue a stay this week, 20 years of court rulings may be about to go out the window Platformer Platformer How Texas' online porn law could shatter a First
Can Ambition Take You To The Moon?
Thursday, May 2, 2024
Yes, with some other ingredients ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏