Import AI 379: FlashAttention-3; Elon's AGI datacenter; distributed training.
Import AI 379: FlashAttention-3; Elon's AGI datacenter; distributed training.If compute isn't everything, why are so many people betting that it is?Welcome to Import AI, a newsletter about AI research. Import AI runs on lattes, ramen, and feedback from readers. If you’d like to support this (and comment on posts!) please subscribe. FlashAttention-3 makes it more efficient to train AI systems: Who else uses FlashAttention: Some notable examples of FlashAttention being used include Google using it within a model that compressed Stable Diffusion to fit on phones (Import AI #327), and ByteDance using FlashAttention2 within its 'MegaScale' 10,000GPU+ model training framework (Import AI #363). Key things that FlashAttention-3 enables:
Why this matters - if AI is a wooden building, FlashAttention-3 is a better nail: Software improvements like FlashAttention-3 are used broadly throughout an AI system as they're used within a fundamental thing you do a lot of (aka, attention operations). Therefore, improvements to technologies like FlashAttention-3 will have a wide-ranging improvement effect on most transformer-based AI systems. "We hope that a faster and more accurate primitive such as attention will unlock new applications in long-context tasks," the researchers write in a paper about FlashAttention-3. Some key points:
Why this matters - why are so many knowledgeable people gazing into the future and seeing something worrying? A lot of people tend to criticize people who work on AI safety as being unrealistic doomers and/or hopeless pessimists. But people like Yoshua Bengio poured their heart and soul into working on neural nets back when everyone thought they were a useless side quest - and now upon seeing the fruits of the labor, it strikes me as very odd that Bengio and Hinton are fearful rather than celebratory. We should take this as a signal to read what they say and take their concern as genuine. What ElecBench tests: The eval tests out LM competencies in six distinct areas:
Results: The researchers test out a few different models, including OpenAI's GPT 3.5 and GPT4, Meta's LLaMa2 models (7B, 13B, 70B) and GAIA models (a class of models designed specifically for power dispatch). In general, the GPT4 models perform very well (unsurprising, given these are far more expensive and sophisticated than the others). Task list for a new AGI:
Things that inspired this story: The fear of death among the mortals; technology rollout philosophies; how many rich people want to ensure their kids don't use much technology; the intersection of powerful AI systems and the physical world. Thanks for reading! You’re currently a free subscriber to Import AI. If you’d like to support Import AI (and fund the lattes which are crucial to its production), upgrade your subscription. |
Older messages
Import AI 378: AI transcendence; Tencent's one billion synthetic personas, Project Naptime
Monday, July 8, 2024
...How the wisdom of the crowd holds true for AI systems as well as people... ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏
Import AI 377: Voice cloning is here; MIRI's policy objective; and a new hard AGI benchmark
Monday, June 17, 2024
Can you evolve your way to Einstein? ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏
Import AI 376: African language test; hyper-detailed image descriptions; 1,000 hours of Meerkats.
Monday, June 10, 2024
Will an open source model get released in 2024 that cost more than $100m to train? ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏
Import AI 375: GPT-2 five years later; decentralized training; new ways of thinking about consciousness and AI
Monday, June 3, 2024
…Are today's AGI obsessives trafficking more in fiction than in fact?... ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏
Import AI 374: China's military AI dataset; platonic AI; brainlike convnets
Monday, June 3, 2024
Plus, a poem about meeting aliens (well, AGI) ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏
You Might Also Like
JSter #227 - Libraries and more
Tuesday, September 17, 2024
With JavaScript, there's always a thing that you don't see coming. I have just a couple of quick things to mention: 1. there's a petition to free JavaScript from its trademark to allow free
New Blogs on ThomasMaurer.ch for 09/17/2024
Tuesday, September 17, 2024
View this email in your browser Thomas Maurer Cloud & Datacenter Update This is the update for blog posts on ThomasMaurer.ch. Remote Desktop Connection (RDP) to Azure Arc-enabled Windows Server
An executive’s guide to implementing generative AI
Tuesday, September 17, 2024
Get a step-by-step guide to generative AI implementation so you can put the technology to work. ㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤㅤ What are you trying to
Even Flow
Monday, September 16, 2024
Brexit 2, Custom Drinks, For-Profit OpenAI, Amazon Bloat, Apple OS Day… Even Flow Brexit 2, Custom Drinks, For-Profit OpenAI, Amazon Bloat, Apple OS Day… By MG Siegler • 16 Sept 2024 View in browser
🕹️ That One Time Apple Made a Console — Analog Computers Are Coming Back
Monday, September 16, 2024
Also: The PlayStation 5 Pro is a Bargain, and More! How-To Geek Logo September 16, 2024 Did You Know Rats, mice, and other rodents communicate not just in the range of sound frequencies humans can hear
[AI Incubator] Fall enrollment is now open 🍁🎓
Monday, September 16, 2024
NEW: We're adding more live coaching sessions
Deepdive – Competitive Analysis
Monday, September 16, 2024
As a Product Manager, staying ahead of the competition isn't just an advantage—it's a necessity.
Daily Coding Problem: Problem #1558 [Easy]
Monday, September 16, 2024
Daily Coding Problem Good morning! Here's your coding interview problem for today. This problem was asked by Twitter. A classroom consists of N students, whose friendships can be represented in an
When Logs and metrics aren't enough: Discovering Modern Observability
Monday, September 16, 2024
Let's return to the previous series and discuss the typical challenge of distributed systems: Observability. We'll continue to use managing a connection pool for database access as an example
The Art of finishing & The browser for research
Monday, September 16, 2024
A new deep dive about a new browser, track everything and understand your life, the story of Figma Sans, and a lot more in this week's issue of Creativerly. Creativerly The Art of finishing &