
Loading summary
A
Today we put out the bat signal and called for an emergency pod because America just experienced an AI Sputnik moment. Kimi K3 released yesterday, shocking the AI world with the largest opweight model ever. And it went straight to number one this week. They didn't just close the gap, they jumped the fence.
B
KME has always been a model that felt a bit different. That's why it was always top of the writing benchmarks. For example, K3 is actually a multimodal model, so it can have all sorts of inputs and it can understand things, which is one of the reasons it's so good at front end.
C
Now it's a free for all between Metta and SpaceX, AI on the American side and now China and Moonshot number three on that Pareto Optimal frontier.
D
Frontier Intelligence is now a totally perishable asset.
C
What are the American frontier labs spending their money on?
A
I think the US government starts a strategy of constraining in some fashion Chinese opwait models from being used in the us.
D
We've had this bunch of the Internet world that information wants to be free, basically. Intelligence also wants to be free.
E
All we need now is some kind of a global. Now that's a moonshot. Ladies and gentlemen,
A
welcome to Moonshots, everyone. The number one podcast in all things AI and exponential. Your front row seat to the coming singularity. Maybe I should say to the singularity which is now.
C
To the present singularity.
A
To the present. To the continuous singularity ongoing. Today we put out the BAT signal and called for an emergency pod because America just experienced an AI Sputnik moment. But more on that in just a moment. Allow me to welcome my magnificent moonshot mates. We have the full quintet with us here today. Alex Weiser, Gross, Dave Blunden, Salim Ismael and Imad Mustaq. I'm Peter Diamandis, your host and abundance provocateur. If your head is spinning at the pace of the singularity, good. Mine is too. And that's the point. Our mission here at Moonshots is to keep you informed, keep you up to speed on exactly what's happening. Most importantly, keep you optimistic with the extraordinary pace of change, the coming age of abundance. Gentlemen, welcome. Thanks for getting up early, wherever you might be. Or IMOD in your case, in the afternoon. I was up at 4am this morning. The benefits of jet lag, but I could have used another hour of sleep.
E
And Ahmad also get up soon after
B
European siesta.
A
Yeah, I've got a workout scheduled right after this. A lot happening, gentlemen, A lot going on. And appreciate everybody's time here. You know, before we get started, I want to personally say thank you to all our subscribers and our viewers. You know, I've had a chance, I don't know if you guys did recently to watch and read the YouTube chat. And all I can say is we love you guys too. You know, our mission here is delivering the news and we spend an ungodly amount of time reviewing, you know, Salim and Alex and Imad, I got your text this morning. Let's add this, let's add that. So, so much going on, I got
E
to say also all the memes of Alex, all the memes of Alex explaining JSpace are awesome. So keep, keep memeing Alex every time you can.
A
Yeah, for sure. And some great appreciation.
C
Figure out, see if you can figure out my J space.
A
Yeah. Well, can we look inside?
E
We'll be able to see.
A
Yeah, we're going to get a readout. And Salim, a lot of love for you on the comments as well.
D
There's some wonderful people out there. You know what's incredible is most YouTube videos are just a kind of a flame throwing festival and ours are completely the opposite. It's really amazing. So, yeah, kudos to you, Peter.
A
Well, no, I just, again, just absolute gratitude and I appreciate the fact that everyone, all of our subscribers and viewers here take the time to listen to the pod. And you know, we're constantly, we spend so much time with our entire team and the entire moonshot mates here just really trying to assess what's going on and deliver it. And we have these emergency pods. So if you haven't subscribed and turned on notifications, please do. Jen. So we jump into the first story. It's a big one.
C
I'll just note that if we do enough of these emergency pods, at some point it turns into moonshots daily.
A
Yeah, or continuous. I still think, you know, moving into an Airbnb together and just turning on the camera.
C
It's going to happen.
A
All right, let's jump in. We've just had a Sputnik a moment that's waking up the US frontier labs like quadruple espresso shot Kimmy K3 released yesterday, shocking the AI world with the largest opweight model ever. And it went straight to number one. A little backstory here. Kimmy K3. And Kimmy is from Moonshots AI, a Chinese lab. And over the last year they've climbed the leaderboard. They put out K2, K2 6, K2 7. Each one closing the gap against anthropic and OpenAI this week they didn't just close the gap, they jumped the fence overnight. They released Kimi K3 and it's a monster 2.8 trillion parameters. And you got to remember the context here. China is doing this while under US export controls intended to starve them of the most advanced Nvidia chips. That's a big deal I want to discuss with you guys. They've completely engineered around the Compute Wall and K3 jumped 17 places to the previous Kimi model, blasting past Claude Fable 5 to land as number one on the front end code arena. K3 is also ranked number one in six other domains. Brand and marketing reference based design, data analytics, consumer products simulations and and content creation. The full model weights are set to drop around July 27, which means anyone on earth will be able to download and run this on their own. Prem gents, how big a deal is this?
C
Alex, I think it's great for competition. Let me first as a preliminary matter, point out some things that have perhaps been slightly less obvious in the coverage. The meltdown, if you will, over K3.
A
It has been a meltdown.
B
Yeah.
C
The first is, as Moonshot points out, they claim in nine of the past 12 months that Kimi models Kimi model series have held state of the art among open weight models. So if that claim is indeed true, over the past year it's been basically Kimi all along. I think that's very interesting. Secondly, taking a look at the published architecture since we haven't actually seen the open weights yet but they're promised later this month, there's no magic in it and that's pretty striking. One can imagine that behind the scenes in Anthropic or OpenAI that they've somehow Sam Altman continues to tease at this that there's some post transformer architecture lurking behind the scenes achieving all of these performance breakthroughs. But taking a look at the published K3 architecture, there's no magic. It's still essentially a transformer. They've made obviously a number of innovations, but well understood innovations concerning how they do mixtures of vectors, experts, how they linearize attention, they have their own special Kimi brand of linearized attention, but it's still basically a recognizable transformer. And I think that the fact that a recognizable transformer like architecture can almost match GPT 5.5 max on the task cost frontier, which we should probably throw up a slide for, I think that's pretty striking. That does raise the question what are the American Frontier labs spending their money on? If you can just use a Transformer to get this close. It's already on the cost frontier, but you can get close like third place on the total. The state of the art for overall AAII performance. What the heck are the American labs spending all of their money on? So I derive great comfort at minimum knowing that the transformer architecture is still alive and kicking.
A
Imad, your analysis here because you've been tracking this. We've been going back and forth on WhatsApp together.
B
Yeah, no, I mean, I think Kimi has been top of various benchmarks. Again, you can pick and choose. And they have had the largest open weight models out of China regularly ever since they almost kicked off a year and a bit ago. I think, as Alex said, the architecture isn't anything super novel. Like there are improvements like the muon scaling that they did with UCLA and kind of other things and they've actually been releasing breadcrumbs of all of these parts. I think what's key here is the underlying data. Kidme's always been a model that felt a bit different. That's why it was always top of the writing benchmarks, for example. And what they've done here seems to be something extraordinary, which is when GLM came out. It's a fantastic model. It wasn't quite up to frontier, but it was text only. K3 is actually a multimodal model, so it can have all sorts of inputs and it can understand things, which is one of the reasons it's so good at front end, although we wouldn't have expected. Again, it's number one in front end versus everyone. And I think this comes to something which I've said before, which is building great solid models is cutting edge manufacturing. Like again, you will have algorithmic improvements and there are all sorts of things coming. But why are Chinese EVs better than Ford's? This actually feels like the same thing, right? Like it's engineering, but it's also like the number one car here in the UK last month was the Jaiku J7 or Temu Land Rover as it's been known. It comes out fully loaded, full spec for like 50k, you know, a third of the price. And this actually feels something very similar. They've known what the ingredients are, the raw materials. They're now putting in an incredibly consumer frame friendly way. And they're just executing that manufacturing process with what they have. Because when you look at the architecture, you look internally, they're still on H8 hundreds, you know, they're like a couple of generations behind on the Nvidia chips, but Then they built it to take advantage of Huawei and Alibaba's next generation chips, which you can see by the static shapes and all sorts of other things as well. And they're just relentlessly going at the engineering and the usability, which is why the front end code, I think, is the one where they're standing out, because they're just like, how can we make it have the most amazing outputs, a personal website to a game to other things. Whereas the US labs are maybe looking in other directions and focusing a little bit on different things.
A
One question real quick is we've always talked about, do we need another breakthrough beyond LLMs to get to AGI? Does this give you comfort that we don't need another breakthrough to really move forward again?
B
It all comes down to definition of AGI, right?
A
Yeah, of course.
C
Don't get me started six years ago, Peter. It was six years ago.
A
But I mean, I guess the question is there's plenty of headroom still to progress these models.
B
Well, I think you have the base model here, right. But then you've got all these amazing harnesses that are coming out and the way that you're using the model to go back on itself. One of the things that's in the Kimi blog post is that it actually designed a chip for itself for its next generation and it designed its own kernels for running as well. And so you move from this model weight to this whole ecosystem that the model itself builds. That feels AGI ish, right. That feels like recursive self improvement. That feels like the ability to learn and adapt new skills dynamically by changing itself. So I think for most definitions of AGI, we probably don't need something new to optimize and make it super efficient. Yeah, there are various ways, even with what we know, it could be more efficient than what we have here. We don't have enough quite compute for it. And new architectures could push us even further.
C
You could say attention is still all you need.
A
I like that. So, Alex, we've thrown up here the, the performance charts and we see Kimike 3 sort of topping the charts in a multitude of places. I don't know if you want to comment on this. And if we could throw up, the
C
AAII scatter plot I think is probably the most constructive one. So this is from the Artificial Analysis Intelligence Index. And this is of all of the charts at this point. This is my favorite one because this one actually shows the cost per task as defined by AAII versus Performance Frontier. So one can sort of Mentally, look at this. For those who can't see it, we see the frontier as sort of a jagged frontier going from lower left to upper right, where in the upper right we see maximum cost per task and maximum overall score is still Fable 5. And then riding the Pareto frontier down and to the left from that we see number two on the frontier is still, as of a few days ago, GPT 5.6 SOL max. And now for the first time, Kimi K3 is number three. It's on the frontier, it's number three both in terms of raw capabilities and also the third point on the optimal cost performance frontier. And I think that's totally striking. We went from a world where, as we mentioned a couple pods ago, where there was this OpenAI anthropic duopoly, to now it's a free for all between meta and SpaceX AI on the American side, joining the upper end of the Pareto frontier. And now China. And Moonshot is now number three on that Pareto optimal frontier. And that's so exciting for any enterprise to the extent it's willing and able to use a Chinese soon to be open weight model to control more of its own destiny. I think this is just such a boon for enterprise sovereignty, it's a boon for competitiveness. We're living in the AI version of for all mankind, where the Soviets landed first on the moon and now the space race never ends. The AI race is now no longer ending with a duopoly. And I think that's a total boon for the future.
A
Light come Amazing Dave, let me pull you in here. What are your thoughts?
E
Well, you know, Peter, you called it a Sputnik moment. If anything, that's an understatement of the implications of this. You know, we had that Alex Karp rant on the podcast last week where he was saying, look, you can't, as a large enterprise, as a government, you can't just throw all of your proprietary weights, your proprietary alpha, all of your intellectual property over the wall to Anthropic and make that the basis of your whole future. But he didn't give you a roadmap to move forward. Here we are just a week later and it's suddenly a free for all, as Alex was saying, a free for all where anyone who reads these weights has the ability to get very close to the frontier and then fine tune for any vertical use case beyond the frontier. And so it gives everybody in the world, every corporation, every government in the world, a way to catch up to the frontier without going through the US AI models. So, you know Sputnik. Yeah, Sputnik times, times infinity, essentially. And the thing I don't like about this particular chart is because the left index goes to 100%, and when you chart it out over the next two years, it looks like an S curve. And so we're in this really steep part of the curve right now, but it implies then we get to 100% and then we've achieved the end. But this, this is actually an exponential where intelligence goes to infinity. So the benchmark saturates, but intelligence itself goes to infinity. And so now it's really, really clear,
A
everybody, you know, the way this works typically is nested S curves. Right. One particular technology tops out, but it builds the next technology that then begins its exponential ascent and so on and so on.
E
Exactly. Exactly. Right. Let me just say one other thing. You know, Alex and I have spent a lot of time working on this Keller Jordan Speedrun. We talk about it a lot. It's a way you take a GPT2 class model, you can find it online very easily. Just look on GitHub, look up Keller Jordan Speedrun. And it's a whole bunch of hackers and AI researchers who are continually trying to take GPT2 way back five years
C
ago in the form of Andrej Karpathy's Nano GPT in particular.
E
Exactly. And try to recreate it faster and cheaper. Faster and cheaper. And if you look at the innovations in that repo, they've been able to cut the original cost of creating GPT2 by 99%. So now it's 1% of the original cost. Yeah. And so everyone doesn't pay attention to it because it's GPT2. And up until today, it wasn't clear whether those same ideas would apply at frontier scale. Now it's really clear that when Elon Musk takes his $16 billion Colossus 2 data center and builds a 10 trillion or 20 trillion parameter model for billions of dollars, there is a 1% cost version of creating effectively the same thing. Nobody knew until Kimike 3 whether that was going to work or not. Now it's really clear that it does work. We're looking at 100x innovations in the software stack, in the kernel optimization, in the mixture of experts. These fundamental breakthroughs that come out of China are giving them 1% cost. So I think Imad gave a great analogy to the car where you can get a virtually identical car for about a third of the price. Here we're talking about less than 1% of the price. To create the equivalent product. Sputnik. Yeah, that's the understatement of the century. This is just. And that's why we're on the emergency pod today.
A
Yeah. Salim, jump in, please.
D
Three points to make. I think it's not so much that Kimmy's beaten, et cetera, whatever. It's the fact that Frontier intelligence is now a totally perishable asset. Like the shelf life is weeks now for anybody that gets to the very edge. And any enterprise or government interested in that very latest cutting Frontier model doesn't have time to actually evaluate it. Do an rfp, look at other models, have a committee internally think about which weather to deploy it. And now you're three generations ahead in the model anyway. So now all the value comes in the architecture that can swap models. And that's going to be the next layer. We call that interfaces in our exo world. That's going to be where all the value resides going forward.
A
Yeah. Amazing. I love this.
C
We need a new term for that. Maybe like the Frontier Liberation Front.
E
Let me say one other thing for the hyper geeks out there. Ahmad said the Muon Optimizer, but he said it very, very quickly. And anyone who's an enthusiast, look that up as well. Because one of the reasons this is happening is because when we built these original models, the very large scale models, we took, you know, 20, 30 trillion tokens from around the Internet, every word ever written by humanity and just dumped it into the training set and said here AI become intelligent given all of this information. But when you look under the covers, the vast majority of that information is Taylor Swift's concert coming up and their wedding like is a whole bunch of stuff. Stuff that doesn't actually drive the intelligence of the model.
D
Significantly the opposite in fact.
E
Yeah, yeah, it's very true. A lot of those tokens actually might slow down the training, not accelerate it. And so purely by pulling out the garbage and stripping down the training set to the relevant subset, it still taxes the model just as much, but it reduces the number of flops, the amount of computation that the model's doing to get to the same level of intelligence. I don't think we're anywhere near done with that problem yet. So you can expect more 10xs to come out of just the Muon Optimizer process and the training data set getting stripped down process.
A
I threw up this tweet from a guy named Alaric that I found fascinating. It's for those not viewing this. It says from anthropic quote fable is an agentic coding super weapon capable of developing cyber and bioweapons at unprecedented speed and scale. We cannot in good faith release it without guardrails. Right. This is the conversation a month ago and China comes back and says laughing my ass off. Here's Fable. But open source, good bleeping luck. So I am curious, how do you guys think about that? That fact that we were so constrained because of the guardrails and here's an open source equivalent of Fable.
D
Well, the Frontier labs have a major problem. They've got three fundamental massive constraints that they can't get around. One is compute and the availability of chips and all the electricity and power that's needed. The second is frontier open source models that are as good as, or in many cases substitutable without much notable difference. And the third is you've got government coming down on you going we need to check before you release anything. I would, I'll make a thumb in the air. Guess the trillion dollars at OpenAI might have been worth shrank by about 50% when the government said we have to review all these models because now it's going to take time to get things out. I think this crashes it by another 50%. I would put the finger in the air value of these Frontier labs at about a quarter of what they were three months ago.
A
If that, I mean if I don't have to spend the money for the API calls and I can just use communicate3 on prem, why would I spend the money? Are they going to be hit by massive reductions in in revenues?
B
Yeah, I think there's a couple of things here. Number one is reduction in revenue. Why do people pay for IBM? You know, why do they pay for non Chinese cars? For mission critical things. I think having us on call entities where you know things aren't going to go wrong will still sustain for a while. So I think revenues will still go up for OpenAI for others. And this is why they built these forward deployed engineering companies as well. And so I think they've still got a way to go. But you know, you have the substitution effect again. This is just like Chinese industrial substitution. Why can't America build industrial things? Why do you have Chinese? Sometimes you buy Chinese, sometimes you buy American and I think we'll see that at least for another year. But then it gets difficult on the cyber attack, security theater kind of things that we've had. You know, I've maintained that we would get to this point. And what does it mean? It means the only form of thing that you can actually do is cyber defense. Like this must be the absolute biggest category in VC right now. Like if you're a talented Stanford MIT grad, build a cyber defense startup that goes into cutting edge and every other company and says, let's use this technology to defend against what's inevitably coming. Because the proliferation of these capabilities is going to increase, but not quite as fast as we think, because what actually happens, and you know, we've done some tests around this, is that GPT 5.6, the cyber version, Fable, etc. Are trained on lots of CVE and cyber data. The Chinese models don't actually have that much of that, so they're not that great. But someone can train that data if they have it into there. And so we'll probably see cyber attack capable open source emerge, I'd say in a quarter or two. So there'll be a bit of a lag there, but definitely for the types of big adversaries, it's going to get a bit crazy.
A
Dave, you want to jump in?
E
Yeah, for sure. I think we glossed over recursive self improvement there. Peter. You asked the question of is this the tipping point? The view of the US government, we always knew it was going to be too late, right? It just moves too slowly. But the view was, look, when we get to a model that's capable of building itself, building the next model, we're not going to let that go out to everybody in the world so they can catch up overnight. Because there's never been a product in the history of manufacturing like, like a car. If you, if you have your state of the art car and you give it to a foreign government, they can't use it to make a better car. But AI doesn't work that way. If you have state of the art AI and you give it to a foreign government, they can use it to actually catch up to you and create state of the art AI. And that became clear to the government, what, a month ago, month and a half ago that Fable 5 was over that line and so they stopped it. But the reality is that Opus 4.8 was over that line and people in China could use Opus 4.8 to create Kimik 3. And so that recursive self improvement line was actually crossed earlier than Fable 5. And that's going to be obvious to the world now because all you need to do is have an AI that's capable of improving its own kernel. It doesn't have to. This is the point I've made on a podcast. Like months ago, people think that RSI is going to trigger when it's Einstein level intelligence, but all it has to be able to do is improve its own kernel and get a 10x step up in speed, which nobody perceives that as being true AGI, but that's all it needs to accelerate itself by 10x. And then the 10x smarter or 10x higher parameter model will be some level of intelligence higher. A lot of people in academia were saying, well, look, we're getting diminishing returns with the parameter count, so a 10x faster model won't natively be 10x smarter. But that turned out to be wrong. We're seeing slowing, but we're not seeing flattening of the intelligence curve. So all the evidence now is that if you boost the raw speed by another 10x, you're going to see genius level AI, and then that genius level AI will boost its speed again. So I think when we look back on this in history, we'll say right around Opus 4.8 was the point where the little spark was enough to ignite a flame. And then a flame can become a fire, and then a fire can, you know, can become a sun. And that's, I think, the way we'll look back on this moment in time. So the cat is definitely out of the bag. The current policy, the current U.S. policy of constraining the next model, there's no way that's going to contain global and corporate proliferation of frontier.
A
Do you think the US, do you think the US government starts a strategy of constraining in some fashion Chinese opweight models from being used in the U.S.
E
well, you know, in two weeks, these weights are supposed to be open weight, open sourced, and then it's, we'll see. Like this probably. If they're, if they're rational at the White House right now, they're spending every minute in a debate on do we negotiate with China immediately and not release those open weights? And I really doubt they'll move quickly enough. I'm sure they'll. Well, I'm not sure. We'll see what happens in two weeks.
A
Fascinating.
D
Can I merge two ideas here?
A
Yeah, of course, Please.
D
You know, Peter, you talked about exponentials and the law of accelerating returns, right? I think it's worth drilling into that because if you connect that to what Dave just said, this is why we've been saying forever and a day on this podcast that this is unstoppable. Ray's original observation was once you have an information based paradigm, you just keep hopping across multiple technologies. So we had vacuum tubes, relays, and then vacuum tubes. In computing, at some point, you can only fit so many vacuum tubes into a room. But that architecture was used to design transistors. Transistors were used to design integrated circuits. And you get these nested S curves. And so what Dave is talking about is as these architectures, all the various pieces of the puzzle get all reinforcing loops inside them. Each of those is like an S curve that starts accelerating the collective. And it's unstoppable. And so it doesn't. There's no limit to where this goes. And this is why people are so kind of freaked out about the upper end limit of this. So important to connect those two dots.
A
Yeah, for sure.
B
If I could just say something. I'm just following on from Dave. So there was an important speech by Xi Jinping a couple of days yesterday. God, time flies. At the World AI Conference in Shanghai, where he basically said, we are going to fully back open source as a public good for humanity. And they're not going to regulate and stop it. This is their plan. It's great for China for a variety of reasons. From the fact they have a billion people whose IQ is about to increase, you know, by having these tools, from the fact they need robots to solve their demographic thing, and the soft power from putting a Chinese educated brain as Tsinghua graduates into every critical system in the world. But they're gonna keep on doing that because they actually have a regulator. And from talking to some of the Chinese labs, it used to take 60 days for a model to be approved. Now it's like a week.
A
Amazing. You know, and Xi Jinping just also announced a regulatory body that they've created, which includes Brazil, different parts of Asia and Africa. I don't know if you guys saw.
C
That means obviously the new belt and road is now focused on AI coming out of China. It's a bizarre future where the Chinese Communist Party is saving American capitalism from itself.
B
It's so true.
D
Let's also note that Yang Zhilin was a CMU graduate. And. And we could have given him a visa to stay.
A
Yeah, we're gonna get to that story in a second. Second, Salima, I'm just. This is an interesting chart here that shows the valuation. So Kimi's valuation or MoonShot's valuation? Moonshot AI valuation is at 20 billion as compared to anthropic at a trillion. And OpenAI basically at a trillion as well. If they were public companies today, I think you would have seen like a 30% stock valuation drop.
C
I'll ask again. What are the American Frontier Labs doing with all of their capital?
A
What are they spending their money on?
E
Well, actually, if you go into the buildings and talk to them and you have any idea at all, they'll give you the capital. They are desperate for more smart people to help because they're trying to deploy and change the world at this insane pace no one's ever experienced before. And they want to deploy that capital much more quickly than they can find smart people who have good ideas to use the capital. But it's a great point. Like, you know, you're sort of saying it in an accusing way, like, what are you guys doing with your capital? But no one in the history of the world has ever had this much money pour into their building this quickly with no prior business experience. We're talking about CEOs that have never run a company before. It's like they're trying. But I mean, seriously, can any human being really rise to the occasion of AI that quickly? So my point there though, is if you're smart and you have good ideas, get into those buildings and propose your ideas. This applies to X Prize too, you know, they are desperate to move that money out the door into something productive that gives them a sustainable barrier to entry.
D
Agreed.
C
And I also think that the Frontier Labs are also asking themselves that question and asking the US regulatory apparatus that question. Anthropic regularly is sending out smoke signals accusing various Chinese Frontier Labs of distillation attacks. And maybe in Anthropic's public mind, that's how the Chinese labs are able to do it, through distilling and capturing reasoning traces. But honestly, like, looking at the K3 performance, I'm not at all convinced that Moonshot is achieving their performance purely or even substantially through distillation attacks on Claude. It just doesn't smell right.
E
No, no, I totally agree. I think, though, that there's a tendency to underweight or undervalue the existence proof. Like just purely the knowledge that a highly scaled transformer running MOE with a Muon optimizer and simplified data, knowing that that works gives you a much more refined roadmap. You don't have to copy, you don't have to cheat, you don't have to steal every trace. You just have to know that that formula works and that cuts your R and D costs by 90, 95%. So I think it's just that simple. There's nothing sneaky or cheaty about it. It's just knowing you're on the right path.
D
I have the greatest value creation idea for ourselves Ever.
A
Okay.
D
Which is we, in nine days, when they drop their open source weights, we release an open source model called Kimi4 under the Moonshots podcast name and IPO, it instantly will be billionaires.
C
So you're saying Celine, what's better than one moonshot? Moonshot's plural.
E
Well, yeah.
D
Why not? Copy the copiers?
E
Let's go.
B
Let's go. Actually, I've got a good analogy for you, Dave. Why do Americans pay more for drugs than everyone else? All the R and D happens in America, you pay the premium, just like tokens, premiums. And then what happens? You have generics elsewhere.
E
Yeah, that's a good analogy. Because that's like a 99% cost cut, which is much more akin to AI than cars. That's a great analogy. You know, Gavin Baker, our friend of our friend Gavin Baker, what wrote a brilliant post, you can find it on X, about the implications of this for businesses. Essentially. Must read. Absolutely. But essentially, all businesses, all stocks other than the foundation AI labs are huge beneficiaries of this. And then like you said earlier, the foundation labs are like, well, what's your future? What's your revenue model? What's. Why are you worth a trillion dollars? I don't quite get it. So you should see a really big reshuffling of valuations in the next week based on that observation. And then any corporation that has its technical act together, you know, there aren't very many of those, but any, you know, if you're a bank, but you happen to be a very good bank with brilliant IT and technical skills, or you have great partners and great vendors, you now have a clear roadmap to controlling your own destiny with your own AI, your own, like, JP Morgan AI. And so I suspect the markets will react to that. If you, you know, put your hand up and say, hey, we have a way to do this internally. We know, we know how to do this with, you know, with our partners, or however you get it done.
D
This is why we call it the organizational singularity.
A
And we're still seeing everybody who's using or trying to use Fable 5 getting downgraded every time they mention biology or mention something that is potentially on the edge. And why would you tolerate that? You know? So in nine days, what do we see? Do we see everybody? I mean, I'm, as soon as it's available, going to upgrade. I'm running Kimi 2.7 on my Mac Studios. I'll upgrade it to Kimi 3. Everybody will. So do we start to see sort of the wholesale US entrepreneurial base of capabilities on K3.
B
Well, I think this, I can give an analogy of this, which is stable diffusion. When we released stable diffusion called four years ago, time flies. You have these really restricted image generators that were a bit better, but they were restricted and they had all sorts of arbitrary restrictions because obviously it's a bit dangerous to have it. You couldn't have likenesses. There was no way to get IP in there, even if it's your own IP and more. And what happened? A hundred million, 200 million downloads and a whole ecosystem that built around that and accelerated generative media, as you said. Why are you going to have this model? Like, I can't even talk about philosophy with it. It downgrades me. Right. Like when you can have the fully open variant of it, even a fraction of the price that you can then customize. A whole ecosystem will build around this and other models. And it has already been doing so. And that's a real danger versus being locked into the single vendor. Which is why I think the labs will go vertically integrated. Like all their customers are now going to be their competition and they're going to be like, okay, I'm going to take you all on.
E
Well, and that directly ties to Mira Moradi and Inkling. Are we going to talk about that story too? That's huge this week.
A
Well, we talked about in the last pod, which was so two days ago.
E
Okay.
A
I mean, and Mira just released Inkling, which is fantastic to see a US open source lab. But the question is, how many more will we get? You know, how many more open source, you know, sort of shocking, you know, Sputnik moments we're going to see. I mean, we have a lot of Chinese labs pursuing beyond just moonshot.
E
Yeah. So Inkling, it's just really telling about where things are going to go because it's, it's designed for you to pick it up as a corporation and fine tune it within your corporate walls to whatever your use case is. So if you're a biotech lab and you're researching and you, you don't want everybody to see your proprietary data, you take Inkling and you tune it internally. But the reason that's telling is because Mira Moradi came from OpenAI. If she didn't believe that pathway was viable, she wouldn't start thinking machines around that thesis. It tells you that the people that are inside the best frontier labs believe that this process can catch up to the frontier. You combine that with Kim EK3 proving it, and it's a different world next week. The other thing that was weird in the market at the end of the week is that things started to reshuffle pretty dramatically toward the end of the week. But in the downdraft, the semiconductor companies also came down, but they're actually going to go the other direction. And this is the point Gavin Baker was making, that this drives up the need for silicon, not down. It changes the whole software landscape tremendously. But silicon is going to be more in demand than ever before and completely sold out, as we know.
A
Can we talk a second about the Nvidia embargo that we put for China? So here, you know, here we see the highest performance models. Was the whole Nvidia sort of regulatory embargo unnecessary? Did it do what we've always done before, which is just spark China's need to develop their own capabilities with Huawei?
C
Of course that's what happened. Of course we did everything that the embargo only incentivized the the Chinese frontier labs to develop and cultivate new efficiencies that, by the way, were always there. To Dave's point earlier about the Nano GPT Speedrun, there's this enormous overhang that isn't fully exploited in terms of leveraging algorithmic and computational and hardware efficiencies to train larger and more capable models. And all these export controls do, I think, is incentivize the Chinese labs, which are already feeling plenty of demand pull to compete with Western frontier models, to leverage those efficiencies, efficiencies sooner. And maybe on balance, although it's superficially bad for the west, now that we've incentivized this new generation of much more efficient Chinese frontier models, in the end, I think it's net good for not just the world, but also for the US to have this fire lit underneath them by Chinese competition that's much more efficient, much more capital efficient, more weight efficient, probably more bit efficient. This is all a net positive. As long as the US in my mind does not set up or fall into some ultimately protectionist regime of trying to prevent what may be construed as Chinese superintelligence dumping on the us that is exactly right. As long as we avoid that, it's great.
E
Exactly what happened. That's exactly right. And I think the US learned a really important lesson in the Vietnam War. And then, because that's over 50 years ago now, it's been forgotten again. And then you have to be reminded. But in the Vietnam War, it was really clear that either you go to war and you win quickly or you Don't. But what you don't do is send in a few troops and then send in a few more and then creep in and like nothing good comes of that at all. The embargo of chips on China was totally harebrained because it was enough to irritate but not enough to actually work. It's just the worst case scenario. And it sparked exactly like Alex said, a huge amount of quantization research which is critically important and under discussed that allows faster performance on cheaper chips. And those innovations don't go away. You know that's going to be around forever. Now let's go to imad.
A
Yeah, imod.
B
I think we've got a. Yeah. Completely self contradictory, but it has some interesting outcomes. So the total amount of compute used for kimik3 is the same as inkling. Wow. And you can tell that because it's the amount of dense weights and roughly we assume about twice the number of tokens trained. Because we don't have it, we're like, but how does that work? Well, you look at their architecture and it's a two and a half times in data to intelligence conversion through the advantages and data mix that they have because they've had to operate in these constraints. And we see that because the first model isn't as good as the second model for inkling, you're going from a trillion parameter model to a 300 billion parameter model about to be released, which is this actually better performance. You see this with the labs and these labs have had to deal with the constraints. But here's something really interesting. I think if you look at that slide that Alex loves and we chuck it up on the screen, what they've had to do is they've had to optimize their inference for Huawei 1910 Ascend chips, for the new Alibaba chips and others. 64 nodes in one. Because this is a big model, like you're gonna have to buy another Mac studio or 2 Peter to serve this. You know it needs like 2 terabytes of RAM. So you see where Kimi K3 is there. That's because they can only use Chinese silicon to run it. They don't have Blackwells, they don't have Vera Rubins. Vera Rubens and Blackwells are designed for these really large models that have really small tinks because it's 50 billion active parameters against 3 trillion total. American companies like Modal, like Fireworks, like Base 10 will be able to serve this model 10 times cheaper than their Chinese competitors because they have access to the Nvidia And AMD big chips. And so like I said, it's a bit ironic whereby the development R and D suddenly has gone there, but there's going to be a 10 to 100 times price drop once this is optimized for the next generation.
E
Vera Rubin well then also Ahmad, you're saying essentially the same thing, but that also unleashes a bunch of chips that aren't currently in circulation, they're underpriced and also a bunch of fabs that can't make a GB300 but they can make an inference time chip that'll run the cheaper Chinese or the lower granularity Chinese model. So a lot of capacity for compute gets unleashed through, through that same process he just described.
D
If I could go up a level and go a little bit. Woo, woo. Right. We've had this mantra in the Internet, we've had this mantra in the Internet world that information wants to be free, right? Basically intelligence also wants to be free. And essentially we've gone over the course of evolution from biological intelligence where you had evolution built in, recursive improvement and then we broke through that to individual intelligence to the person of a species to collective intelligence like markets or networks. Now we have AI which can scan across all the data to create a whole other level of intelligence. So this is not stoppable. And so any entity or domain or government or whatever that tries to constrain it always, always, always, always fails. And so it's just a fundamental law of nature that you cannot constrain this and it's just not possible. Why people bother is what really blows my mind. It's a very scarcity mindset to try and think about it this way. The faster we get to better intelligence, the faster we get to abundance, the faster we don't need to fight over anything.
C
I can't disagree with you, Salim. I wrote an entire paper on arguing intelligence manifests in the physical world as maximizing future freedom of action to so here's to the Frontier Liberation Front.
E
Well, you don't want to. You don't want to.
D
Now you remind me of the Monty Python thing where there's the popular People's Front and the people's Popular Front of Judea.
C
We need T shirts.
F
This episode is brought to you by Blitzy Autonomous software development with infinite code context. Blitzi uses thousands of specialized AI agents that think for hours to understand enterprise scale code bases with millions of lines of code. Engineers start every development sprint with the Blizzi platform bringing in their development requirements. The Blitzi platform provides a plan, then generates and precompiles code for each task. Blitzi delivers 80% or more of the development work autonomously, while providing a guide for the final 20% of human development work required to complete the sprint. Enterprises are achieving a 5x engineering velocity increase when incorporating Blitzi as their pre IDE development tool, pairing it with their coding copilot of choice to bring an AI native SDLC into their org. Ready to 5x your engineering velocity? Visit blitsi.com to schedule a demo and start building with Blitzi today.
A
Let me bring up a related subject to this story here that I have a pet peeve about. And it's, it's this one. So you know the founder and CEO behind Moonshot AI, Yang Zhilin, you know, didn't learn his craft in Beijing. You know, he earned his PhD at Carnegie Mellon, you know, one of the best computer science programs in the world, in Pittsburgh. And we basically trained him up, we admitted him, we trained him up at one of our best institutions. And then when he gets his PhD, you know, he doesn't get a green card, he goes through the hassles of trying to get a visa and he goes back to China and he builds Moonshot AI there. Just a moment and talk about. You know, I've stated publicly so many times that I think when anybody gets a PhD, they should get a green card stapled to the back of it. Why are we sending the most brilliant people who come here to get educated back home, whether it's to China, whether it's to India, whether it's to Brazil. Why don't we enable them to stay here and build? Gentlemen, comments on that.
C
Okay, so I did some research on this and I think the story is not what it seems to be. So a little bit of chronology first. So Yang Zhilin, according to my research, he starts his PhD after undergrad in China, starts his PhD at CMU in 2025. Fall of 20. Oh, sorry, fall of 2015. Okay. Then approximately one year later, he founds a startup while a PhD student at CMU. The startup is named Recurrent AI. Where is Recurrent AI based? It's based in China, it's not based in the US. So one year into his PhD program, he starts a Chinese AI startup while still doing his PhD at CMU. That's interesting. And that's a problem. This also runs counter this sort of a narrative violation for, oh, we wouldn't staple his visa or whatever. And then he goes back to China. No, actually one year into his American PhD program, he starts a Chinese AI startup, then he graduates in 2019 is my understanding. My understanding is he had offers from Google, Facebook, Huawei and others upon graduation in 2019. But he goes back to China because that's where his startup Recurrent AI was actually incorporated a few years earlier. And so I don't think necessarily this is the case where either the US was unwilling to retain him or even President Trump somehow, through some policy, was driving away this particularly talented Chinese graduate. He started his company during the tail end of President Obama's term in China.
A
And Alex, I appreciate the deeper dive that you did. Thank you for that. The point still stands. And you know Salim, you and I have seen this so many times, right, at Singularity University. Dave, you may have seen this at mit. I mean the fact of the matter is a lot of the most brilliant students aren't given the opportunity to stay and develop here. Dave, or actually Imad, what are your thoughts on that being someone not in the us?
B
So if I can just give my two cents. Yeah, I completely agree with it. And here's this crazy thing. The math and the numbers are all there. What is the value of a PhD staying in America? It's actually quantifiable and there've been multiple studies on that. Dave does a great job obviously of converting them into startups, into innovation. And then there's the other thing that shoots in the foot, which is American companies can't invest in Chinese companies because of regulations and other things as well. Some of them are Chinese. But look at the trouble that benchmark got in for investing in Manus, for example. So I think it's kind of twofold. But I completely agree that if you've created or contributed to creating a valuable asset. Most foreigners stay in America after they do their PhDs, but too many don't have a very direct path, despite the math proving that they will add value to the American economy.
A
Yeah, I mean another point just to make here is, you know, the AI race isn't only about chips and compute, it is about people. You know, key people are still driving the greatest value, at least for the moment.
C
I would argue not just people, but also to my earlier point, it's about where the startups get domiciled. There's an alternative world where he, through whatever immigration oriented regs, was deterred from starting his first AI startup in China while still an American PhD and we incentivized him to start recurrent here in the us And I think there's maybe an alternative counterfactual world where recurrent was American. And then its arguably intellectual successor, which is Moonshot AI, also remained domiciled in the us and then he followed his own startup to stay here.
E
Yeah, one thing that came out of the story is when people come from India to get educated in the US, they overwhelmingly stay. When people come from China, about 80% of the time they go back and it's just a difference in the local economy. Going back to India to start your company as a non starter, it's just so unlikely to catch. But going back to China, it's a thriving ecosystem, lots of support. So going back to China to start your company is actually not a bad plan for a lot of people. I didn't realize that until this report came out. But as Peter was alluding to, we have tons of friends from, from MIT that came from China and I don't want to put them all in one bucket because there's a really clear distinction to me between people from China that are Hong Kong, Taiwan, whatever, that come over that don't really align with the Chinese Communist Party at all. The fact they kind of hate it. And then you've got Chinese people that come over for an education. And in one case at bu, a very good friend of ours is the dean of Computer science at bu. And there was this, this massive crisis because there's a concerted effort by the CCP to plant specific students into BU to gather specific knowledge and they were given tasks. You have to go study this, learn it and then send it back. They didn't know what to do at bu. It's like these are effectively trained spies that got into our PhD program, but we weren't ready for it. What are we supposed to do? They want to be highly ethical, so they don't want to just dismiss the students. So I don't know how they resolved that. So you got this really like, that's a very different thing from the bulk of Chinese students who are, you know, they don't align with the CCP and they just want to thrive in the world and they're happy to start their company here or anywhere else.
A
And don't forget when we looked at the Frontier labs, I mean originally in the early days of Xai, for example, and in Meta, like 50% of their research staff of their research PhDs were Chinese Americans and extraordinary. And the Chinese, every year in the Math Olympiad are at the top of the scoreboard. There is an incredible wealth of capability here that I think most companies desire to retain inside. Salim, you were going to say.
D
Yeah, two things here. One Is, you know, the asymmetry of the talent, I think, is the really important part here. I made this point a couple of podcasts ago. 70% of the elite AI researchers are not US citizens. They're in order, Chinese, Indian, Taiwanese, and UK. And so that's a huge problem. Stapling in a green card is the easiest thing we could do with zero friction to then give them incentive to stay here and build here. The US has massive asymmetric advantage for the rest of the world. It was better to build here than anywhere else in the world. And that's starting to become less true. And that's why people are going back to China, increasingly back to India, even to do things, despite the kind of the friction that exists. Trying to do something in India. That is the part. The failure of the US to fix immigration is one of the biggest problems this country has right now.
A
Amen. All right, I'm going to move us forward here. I just want to put up this slide. You know, since mid April, we've seen 13 new Frontier models launched, an average of one every 10 days. Just, you know, comparing this to 2025, we had eight Frontier releases over the course of a year, one every 50 days a year earlier, in 2024, we had six releases, one every 60 days. And it doesn't seem to be slowing down. And then, Imad, you sent me this morning this tweet from Elon. Thank you. I'll put it up here. This is Elon's tweet. Our 2 trillion model, which is better than our 1.5 trillion in every way, will finish initial training next week. It might be able to exceed Kimmy, but with speed and token efficiency close to our 1.5 trillion, aka Grok 4.5. So, I mean, this is the number one piece of evidence that we're living in the singularity. The speed at which this intelligence is accelerating is insane.
C
Peter, it gets better. If you take Suhail's list of frontier models and the dates, and you regress an exponential curve to the predicted frequency or time period between model releases, which I did just as an exercise, you find that at the present rate, we're going to get to daily frontier model releases by. Wait for it. January.
A
By January.
C
By January, we're going to see daily New Frontier model releases if this exponential trend continues, which basically implies continuous versioning.
A
So I guess the question is, what does that really mean, right? What does it mean to have a new release if it's a continuous process?
C
I mean, maybe it means that we'll have to do our daily moonshots episodes about something other than point releases from the frontier labs. We'll need something new to talk about because it'll just be updated behind the background
A
like my son Jet said. Okay, so another release a little bit better. I mean, like dad, come on, like, what's reeling you here?
E
Well, actually, yeah, we'll see later in the pod some use case demos, but I think those will take over because it's much more exciting when you see a tick up in the intelligence, you're like, yeah, so what? Like, look what it made. That's what really gets people's attention.
A
But let's take a second and just look at that because I skipped over it. But I think one of the things that's interesting here and I'll just play these is what we're seeing is recreate your favorite game. And on the right hand side of the equation here, we're seeing a web browser, a browser based web app simulating an Apple desktop. I think this is, you know, we haven't talked about what the implication to the gaming industry, which is huge. Right? I'll stop that.
E
I think it took over your computer there. It won't stop now.
A
But I mean, what was Fun the last 24 hours was seeing everybody sort of show their use of Kimmy K3. It's impressive. Everybody becomes a creator, everybody becomes a maker.
D
One warning though, One warning, one shotting a game is very different from building the entire ecosystem and the customer service and the marketing that goes around with it, etc. Etc. So you really have to be passionate about that domain. But the friction of getting a game launched per your personal interest or your fascinations or your a particular type of game that you want is near zero now and that becomes really interesting.
E
Yeah, yeah. It's mentally taxing because if you take it to the limit, which is very soon, I can one shot prompt to create anything and then you're sitting with your corporate exec staff saying, well, what do we want? Well, we've never had the ability before. We never really think this way. So then you have to kind of stretch your brain to like, well, what's the purpose of our organization in the first place? Yeah.
C
Okay, here's a thought. Historically on this pod, we've done calls to action to submit outro music videos. What about a call to action to submit an outro video game that people have just casually created?
E
That's cool.
A
So going to your point, Dave, I think having taste, having imagination, understanding what the public wants, I think these become the Scarce elements. And you know, for entrepreneurs out there, as you're seeing this capability, I think the entrepreneurial mindset and the ability to imagine something even greater. I mean, what happens when you're unleashed in what you can make? Yeah.
E
And visualizing happiness is, you know, we're not used to trying but you know, a lot of people don't manage their own happiness particularly well because they, you know, they have to suffer through their daily job, they have to suffer through whatever, you know, mosquito bites and geography. Just you have no choice. Given choice, what would you do? And that's so liberating for the mind. But because we're not used to thinking that way, we're not ready for like there must be an infinite number of things. Like the one that's easy for everybody is medicine and biotech. Like at least I want to be healthy. You know, that's an obvious one. But what about all the other things that make humanity happy? Have we really thought through what we could voice, you know, prompt tonight.
A
And it's really like we are godlike in our abilities.
E
Hence the name of your book.
A
Yeah, well, I mean it's. The idea is being, being a, being a creator or a maker. Right. Versus a consumer or a taker. Salim, I saw your eyebrows go up.
D
No, I'm, I'm just agreeing with all of this. I think this is, there's. It's such a magical time to be alive. Everybody listening to this podcast, but please think up some business idea, project impact, project, whatever and use AI to go build it.
A
I just want to add maybe just
C
briefly, Nick Bostrom speaks about this a bit in Deep Utopia. Peter, you and I speak about this quite a bit in solve everything. I'll just outright suggest folks, if listening, speaking just for myself, I'd love to see an outro video game that you casually create. Maybe something in the theme of the moonshots pod. Since evidently we've about talked completely solved and cooked music video creation.
A
A first shooter game where we get to take aim at awg. Oh no, please, no, no, no.
C
Ideally a non violent outro.
D
Be careful.
B
Non violent Civilization Tech Tree. That's what you need.
C
Civilization Tech Tree game would be great.
A
I want to hit one point. You know, everybody watching and listening here. You have two options when you hear about this extraordinary ascent of Kimmy K3. Fear might be one and the other might be oh my God, what an extraordinary time to be alive. Right? Hope and excitement and just an abundance mindset. And rather than fear, realize you are being unleashed your creativity, your ability to do whatever you want. The ability to create passion, your purpose. Find what Salim and I talk about so much is finding your massive transformative purpose. Just to distinguish between two. A passion is something you love doing. A purpose is something you love doing that actually benefits the world. And so if you can connect with that and realize that you can without any background, I mean, I think this is one of the most important things. You don't have to be a computer scientist, you don't have to be an expert. You have to be purpose driven. And if you use these tools, you can make a dent in the universe. You can improve humanity at an awesome scale. And that's what entrepreneurship is.
D
Three steps. Read Alex and Peter's paper. Solve everything. Pick the biggest problem you dare to pick. Go download the Organizational singularity Claude skill, which is free and start building.
A
Awesome.
E
Yeah. Note for our production team too. Here. It's so cheap and easy now to do things like Alex suggested. Make a video game. We should collect and post some examples for the audience so that they can say, oh, that's what Alex was talking about. But it's, you know, just a little roadmap is all people need. It can be this long if the
C
audience doesn't send in amazing moonshots oriented video games as outros. I promise I will create a cyberpunk FPS but it'll be a non violent FPS if you can imagine that. Oriented around moonshot.
E
What are you shooting at?
A
You'll be tickling bunny rabbits.
C
No, no, no. Okay, fine. It'll be a cy. It'll be a cyberpunk FPS where you're where we're cooking every problem. How about that?
B
It's a first person solver, not a first person solver. Love it.
A
Oh God.
E
You got to do the tickling bunny rabbits too though.
C
Okay, I'll tickle bunny rabbits.
A
Fine. I'm going to move us to our next story.
D
Please let's stick with the music videos.
A
They're great and imod. This is one you sent over the transom that I added here. So if Kimmy K3 is the frontier going big. Trillions of parameters in a data center. This story is about frontiers going small. Small enough to fit on your smartphone. So Bonsai 27B for billion is the work of Prism ML. It's a US based AI startup out of Caltech. It's run by Babak Sabi, backed by Khosla Ventures, Cerebrus and Google. It's the first 27 billion parameter class model to run entirely on a smartphone. Not A stripped down version. It is built on Qin 36 27B. Imad, tell us about this. Why is it important? We just talked about small language models with liquid AI on our last pod.
B
Yeah, this is one of Dave's favorite topics, quantization. Right. And Prism ML and actually Tencent, which I'll talk about in a second, have had massive advances in being able to take model that's been trained in a 16 bit architecture or an 8 bit architecture like Kimi is basically 8 bit 4 bit and take it down to ternary which is three bits of information or binary.
C
So ternary is three values. One approximately 1.58. Yeah.
B
Bits 1.56. Yeah, like so three values. You geeks.
E
This is really, really important. Pay close attention attention geeks, because this is a really important topic.
B
Well, this is again the accuracy thing. So what Prismml managed to do is they managed to get the model down to ternary, which basically means. Well, I think it was 6 gigabytes for the model. This 27B model, which is really performant. I think it's basically GPT5 class from memory.
D
Wow. Wow.
B
With a 5% drop in accuracy, they managed to get it down to 6 gigabytes. And with a 15% drop in accuracy, accuracy down to 4 gigabytes.
E
And you can get the accuracy back too by expanding the size of the network a little bit. Sorry.
B
There's various things you can do. And so this is a big deal because it means you have 1001020 IQ buddy, that can work on your smartphone
A
again without an Internet connection.
B
Without an Internet, it's smaller than a video game.
A
You literally have this level of intelligence in your pocket all the time.
D
Is it live?
B
It's live. You can download the weights right now. You can run it on your smartphone exactly the way you go. But this is the super interesting thing. When you reduce the bits, it also increases the speed. So from 16 bits down to 3 bits it's a 5 times improvement in the speed. And there's another article or another release which is Tencent's latest model. This is actually the old WizardLM team who had to leave Microsoft because Microsoft wouldn't give them computers is very ironic. They managed to get binary compression so taking it all the way down for their High3 model. So it's now the best on a GGX Spark or a Big MacBook to take a 300 billion parameter model the size of the new model that's coming out of inkling to work on binary with a 5% drop in performance.
F
Wow.
E
Yeah. And now I'm gonna. Sorry, Ahmad. This is so important.
B
Yeah.
E
I'm only going to say this once on the POD because we're investing a lot of. Lot of companies that are working on exactly this. I don't want to tip it too much, but the implications of what Ahmad just said have. There's one more step there, which is if I imagine like all of this intelligence under the covers, the computation going on has always been matrix multiplications. So I have a number, I multiply it by another number, and then I add two of those together. It's called a Mac, a multiply accumulate. If one of those numbers is just 1, 0 or minus 1, I think we can multiply a number by 1, 0 or minus 1 pretty damn efficiently. That's the efficiency Iman's talking about. But it also opens the door for new ways to compute. What Alex has been saying for a while, if anyone listens through it, we're going to discover new physics, but also new substrates on which we can compute. And we're going to discover that computation is possible virtually anywhere, in crystals, in liquids. But the computation we're looking for is simply 10 minus 1. So it really narrows the focus on where we look for these computing substrates. That'll take AI to ask. So what we're envisioning right now in the Dyson Swarm is a bunch of GPUs from Nvidia sitting in a satellite with a solar panel and a radiator that's only going to last a couple of years. Something very different is going into space. Something much more like, you know, Star Trek with, with crystals and holographs and things that are capable of doing the exact same mud.
A
It's a linear here first, guys.
D
I mean, this allows.
A
How, how efficient does this get? How compressed does this go?
E
Oh, my God,
C
Peter, on that. So, so this is something I think about quite a bit. So the most quantized Banzai model that we were just talking about, I think is approximately 1 and an eighth 1.125 effective bits per weight. But you could ask the question, like, is one bit per weight the limit? And the answer is no. We can go below one effective bit per weight. How do we do that? We do that with sparsity and quantization and low rank factorization. And by the way, that's what we're starting to see from some of the labs, in particular, like Samsung, for obvious reasons, Samsung wants to be able to host highly capable frontier class models on their own edge devices like smartphones, just in the Past two months, Samsung published a model called Nanoquant that breaks the one effective bit per weight barrier. So it's sub one bit which I think we're going to be talking quite a bit more about in the future using a variety of tools. And so this is my extrapolation episode. I went through the exercise of extrapolating frontier quantization out and naive extrapolation finds that sub 1 bit quantization is going to go mainstream sometime in the next year.
E
And then overwhelmingly likely photonic, the speed of light in photonic will be the way we're computing in the future.
D
Imad, you made a comment. Imad, you made a tweet a couple of weeks ago that said we're gonna get fable level capability running on a normal MacBook in 18 months. Right. This is essentially the path you're talking about.
B
Yeah.
D
So sorry, please go ahead.
B
Yeah, so if you look at what Nvidia did with their last Nematron series, they took the big model and they actually distilled it with logic as they're called, down to a smaller dense model. What's the difference between a 27 billion parameter dense model and these really big sparse ones? When you have the model weights, you can actually do proper distillation, which is a bit different from the reasoning traces. And so we're going to see is models like Kimi get distilled down to perfect data sets for smaller models that will be trained four bit and then cast down to ternary or binary or even lower in terms of the bit weights. And when you actually look at like look at Quen Max vs Quen 00:27B, you can actually extrapolate what the sizes of these models will be. As you move dense and you go through the whole process, you end up with a model that works on 16 gigabytes of RAM by the end of next year. That is the level of Kimik 3. And you can even extract all the knowledge out of Kimik 3 because it will be open source.
D
Wow. So that means every vehicle, every robot, every manufacturing, every device in the world has their own built in persistent intelligence and can make autonomous decisions at the edge for whatever task they can. So this decentralizes capability at the most infinite level.
B
And it could go even one step further when you get down to ternary or binary. Actually ternary is better for many things. You could build custom photonic silicon or even etch onto the silicon itself. The zero is just it doesn't have a path on it. So you can etch the model weights once they're good enough. And that leads to an actual increase in the total speed and you don't need to use the smaller silicon anymore. So the cost of intelligence is going to drop by 100 times anyway by the end of next year, just due
A
to the new chipset speed running Star Trek.
C
And I think this is what imagine Dave was talking about earlier, that as we move potentially to ternary or even sub one bit, it's far more ergonomic to adopt post CMOS type architectures underneath. There's plenty more room at the bottom.
B
Yeah, my bet there is. 0.78 will be the bottom. So I'm going to put that as a marker today.
E
Okay. We can do our end of year predictions on that one. That's a really very specific number.
D
I mean, okay, I'm going to be like four times over to just kind
A
of figure out, do you want to
C
go around quickly and ask everyone like, what their their favorite quantization end game is? Ahmad, it sounds like you have a bizarrely specific one.
B
I'll post the details of that soon. We'll let everyone else have a think about it first and then a future episode.
A
Okay, Dave, you were going to say, Dave?
E
Oh, I was going to say that the most likely forecast based on everything Emad and Alex just said, we're expecting 100 to 10,000x within three years on just the raw compute through quantization and new compute methods. And that's, you know, that's multiplicative with the other algorithmic improvements. It's really hard to forecast. So, you know, realistically, a million X.
A
Let's pause there one second, Dave. And just for folks to absorb that for a moment, that, you know, we've seen this incredible speed in performance and intelligence and we're about to see what is 10,000 or you know, ad algorithmic improvements get you to a million. What does that feel like over the course of what, the next three years? I mean.
E
Yeah, three years. Yeah. Well, one thing it feels like for sure is that the AI is doing things that you really desperately want, but when it explains to you what it did, you just can't keep up. I'm already feeling this with Fable 5. You know, I've got so many Fable 5 agents running and they're doing the outcomes are exactly what I want, but it's like, well, what did you do? And I can't get through it all.
A
I had this conversation with Ray, the point at which AI is asking and answering questions that you can't even grasp.
E
Yeah, no, that's very Soon. So to dive back to our Kapara conversation. The idea of slowing it down is nutty. There's no regulatory concept of slowing it down that makes any sense. All we need now is some kind of a global inspection and global partnership to monitor it and then just take advantage of all the abundance that's going to come from it, all the new medicines, all the new capabilities, all the global happiness. It's imminent. We just need to unleash it, don't slow it down, but inspect everything. This whole mechanistic interpretability is going to become the most important thing that anyone can work on. And we just need global transparency and full throttle.
A
I think this is one of the most important podcasts we've ever had, guys.
E
Mind boggling Sputnik moment, you call it.
A
Sputnik moment.
C
Moonshot's brought to you by Moonshot,
A
our new sponsor. Yes. All right, I'm gonna move us along to another fun story, one that I love talking about. It's called predicting the Future. So there's a guy named Philip Tetlock, he's a political psychologist at University of Pennsylvania who authored a book called super the Art and Science of Prediction after he identified what he called a group of superforecasters. These are ordinary folks who through disciplined reasoning consistently out predict even CIA analysts with classified information. He scores this on what's called a Breyer score, where lower is better. So now the benchmark that pits AI against these superforecasters is called the forecast bench. And it's been tracking a steady year long climb as models close the gap. We've talked about this before on the, on the pod. Well, the newest numbers have just come in and according to the Forecasting Research Institute, for the first time, several AI models are now statistically indistinguishable from four from superforecasters. So the implications, you know, if an AI can forecast novel events at superforecaster level, then every decision that we make in insurance, investing, policy, geopolitics, corporate strategy gets a cheap, tireless, superhuman advisor always on. I find this fascinating. The data is out there and the ability for an AI to gather it and make predictions. So at the end of the day, every political decision is going to be modeled this way, every investing decision is going to be modeled this way, and this becomes sort of the differentiator. So who wants to jump in on this one?
C
I'll jump in. I absolutely love this to pieces. First, a few additional pieces of context. So the number one AI superforecaster is from a British startup named Cassie Short. For Cassandra, who of course made predictions but wasn't listened to. Interesting is founded by a British intelligence officer who served in Afghanistan and advised the British government and then formed this in part inspired by superforecasters. What I think is really interesting, though we've spoken when we've talked about these sorts of stories in the past, about Isaac Asimov's psycho history and other rifts. I want to try a new riff here, which is an interesting thought experiment. What happens when hyper forecasting is not just super forecasting? Hyper forecasting is connected to capital markets. What happens when the AIs which are already AI algo traders are already completely dominating by volume public securities markets? What happens when they have better internal autoregressive models of humanity than humanity does of itself in some sense, in the same sense in which large language models were trained off of the autoregressive task of predicting the next token of Internet text better than humans can, and now LLMs can predict, at least from a perplexity perspective, the next token. I'm going to say in this sentence, probably faster than I can generate it myself, what happens when these hyper forecasters are able to generate the next actions by humanity collectively faster than humanity can take it? That's sort of the ultimate market squeeze efficiency outcome, where literally I think capital markets will be, where this is maximally interesting, where the prediction is actually preemptively shaping the action of the market. And I think those who were so dismissive of the efficient market hypothesis, I think the EMH is going to be crowned king of the capital markets once hyper forecasters like this are ultimately plugged in, which seemingly is imminent.
A
I think this leads to wisdom, right? I think this is one of the most important things. And I've written a substack on this, I've talked about it in the past. And if you think about when you go to a wisdom council and you ask what should I do? You go to that wisdom council because they've had so many experiences in life, they can tell you go this path, it's not going to succeed. Go down this path, you have a higher probability. So imagine a world in which everything's being simulated to the point where an AI can tell you what is the maximal path to take for world peace, or to find your spouse, or to determine how to answer to your kids. If you can literally simulate society on a level, we have a godlike support structure to help us navigate the decades ahead.
D
Well, if I make this practical to an organizational level, right? Think about most high level management Capabilities like budgeting, hiring, decision, product launches, investments in various things. Each of those is essentially a forecast, right? But you never record the probability of that or score the accuracy of that. Once you have AI forecasting that approaches that capability, this means senior management essentially evaporates because most senior management is there because they have deep expertise. If you're the head of supply chain for BMWs because you ran supply chain for Spain or you ran supply chain for that engine over decades you built up your experience that to manage that domain. Once that judgment and hard to quantify. And now an AI system essentially can reproduce that without your biases that you have in that are inevitable for human systems, really important that essentially wipes out all senior management expertise. So now you need to focus even more on purpose and what you're trying to accomplish and the objectives you have, etc. It completely changes the game for senior management in any company and any government.
A
Yeah, Imad, yeah.
B
So you know, it's a topic close to my heart. In my bestselling book, the Last Economy, I actually describe how the mathematics of generative AI can apply to economics. And soon we'll have a paper coming out that derives all of economics from the same math of generative AI. Every single equation. It's kind of crazy, but one of the nice things here, even the incorrect ones, even the incorrect ones, it shows them as limits and why they're incorrect, which is fantastic. But one of the interesting things in psychohistory in Isaac Asimov's foundation, he says that entire groups and populations can be modeled like gas. And the equations of gas are the equations of diffusion models which turn out to be better than humans at prediction. And we are going to release a whole bunch of studies around that on economic prediction where they're outputting, performing. But then this raises something very interesting. You know, Peter, you said the wisdom, you know, Salim, you've said no senior management. The way these models will start entering is second opinions, medicine, business policy. But then the liability profile is going to go crazy.
A
Matchmaking, matchmaking.
B
Well, matchmaking, yeah. We have some dark things there like black mirror and other things. But think about it this way. If you make a decision, decision not approved by Dr. AI, your insurance premium goes up like that. You know, if you take drive and you don't drive, according to fsd, in a few generations your insurance premiums go up like that. And that recursion is something that's super interesting because in foundation you had three requirements for psychohistory to hold, one of which was that the population is sufficiently Large. And that could be like driving a car or entire economy economies. The next thing is lack of technological advances of sufficient levels or technological stagnation because that can change the entire landscape of what's new. And the final thing was ignorance. And so, you know, Alex just mentioned these things coming into the market change it. But these things coming into making a healthcare decision or a government decision or a company decision actually changes the way it's like, hey, you're my match made in heaven according to the AI. How can you argue against the AI? Worst pickup line ever right now. But who knows in a few years, right?
C
This is very meta imad. The sort of reflexivity in economics, I think many would call it. If the best predictor ends up being named after Cassandra and no one believes
B
it, they can make the money. It's okay.
C
You can slice the irony with a knife.
A
Dave, have you seen any startups in this area?
E
No. Shockingly no. And you know, Safe Superintelligence Ilya Sutskever may be doing a version of this, but they're keeping it in house, you know, and launching it toward markets and printing money internally. But the version I'd love to see very soon, I think a huge amount of human unhappiness comes from consumerism and consumer marketing. And you know, like Homer Simpson comes home at 6pm, cracks open a beer, lies down on the couch and starts channel surfing. And then, you know, like naked and afraid is on, ends up watching it until falls asleep on the couch, wakes up the next morning with a hangover, having not brushed his teeth, kicks the dog and ends up with unhappy kids. Like that chain of decisions is so bad. But there's no explicit decision to live that life in that chain. Right? It's just you just reacted to the beer ad and then you went down this chain. And I think AI is going to be an incredible coach to say, hey, dude, what if you take this alternate path and here's the outcome, you're going to get to love that. That to me is forecasting used correctly for just changing. Like, are we anywhere near optimal? And the answer is no. If you objectively look at your life, nobody's near optimal. But with a little AI assistance, you can get on a much better path. But what we do right now is it's massive consumerism. You're reacting to billboards, you're reacting to TV ads. It's telling you you think you need certain things and people tend to get sucked into these pathways. I think we can get out of those pathways with AI.
A
That's Brilliant. Just to say, first of all, there is a rumor out there that Ilia SSI is going to be releasing something very shortly. I think everybody's feeling the pressure to release. We saw that with MIRA coming out. So interesting to see. But the point you made, I think is brilliant is are these labs actually pulling their punches holding on this capability to generate revenue on their own? I mean, if you had this super forecasting capability in the markets today, you would do that. I remember having a conversation with Eric Schmidt who said, listen, if Google wanted to maximize its income, it knows exactly which companies are going to have a stock bump in the fourth quarter because everybody's googling this product or that product. We have advanced information about where the sales are going to be and which products are going to peak. But if we could only do that once and then we'd be shut down. So interesting to see if these companies. And Alex, you and I have talked about the fact in Solve Everything, the notion that the greatest money, the greatest income these frontier labs are going to make is going to be as they solve scientific breakthroughs and superconducting and age reversal and so forth.
C
Exactly. And maybe just a footnote on the Google story. So I've had this conversation with Google execs many, many times over the years. Totally agree with the premise that if Google were to attempt stock trading based on arguably insider or unfiltered insider information passing through the query stream, that's a one and done type shutdown scenario. But there are other things that Google hypothetically could be trading besides public securities that would necessarily have the blowback. For example, again, hypothetically, foreign exchange rates.
B
Yeah, and I think that you have to be careful here though. Like, I think there's the, there's the market side of things and you know, like, maybe I will launch a hedge fund based on our own stuff. But there's the moral side of things. Maybe not okay, of course, at any rates. But look, there's the moral side of these things. Like, it's fantastic that we can optimize ourselves, but who controls these models and the advice they give can control vast weights of humanity. And there needs to be a real discussion about this because it's like the people that follow their GPS into a lake.
E
You know, that's the risk.
B
We're going to rely on these far too much. And again, how can you debate it in just a few years time? Like again, you will, it'll be more expensive not to do this. You will be penalized for not listening. And if we're all Watched over by machines of loving grace. We need to know whose grace that is. And again, that discussion is. Start now.
A
Yeah.
E
Salim just said beer.
D
Homeward drinking beer, advised by AI was not on my bingo card for this episode. That's all I'm gonna say.
A
Welcome to the health section of Moonshots, brought to you by Fountain Life. You know, my mission is to help you use the latest technologies, including AI, to not just do your work at home, teach your kids, but to help you live a long and healthy life. I'm here today with an extraordinary physician. The chief medical officer of fountain life, Dr. Don Musailam. Dawn, let's talk about cancer. You know, I know from the member database that we have at Fountain, our members who come in who think they're healthy, it turns out 3.3% of them have a cancer in their body they don't know about.
F
That's right. You know, the majority of cancers that we screen for, those aren't the ones that are necessarily taking the lives when found at a late stage. We know that when cancer is found early, the chances of for cure are much higher. We know it's much easier to treat a cancer when found early versus when found late. What we're finding in our members is over 3.3% were found to have these cancers that were otherwise wouldn't have been found or detected.
A
Yeah, you know, it's interesting, people, you don't feel the cancer until stage three or stage four. And if you don't know what's going on inside your body, it's like driving your car with your eyes closed, and you can know. And so when members come through Fountain, how do they detect cancers?
F
So we're doing full body mri, and we also do early cancer detection screening. This is very, very important. And these are not typical tools used in the conventional care setting. When it comes to prevention. This is a hard thing because currently, these are not studies that insurance would yet be covering. But the goal is to collect these numbers, do the research, and work hard to democratize wellness.
A
Yeah. So at the end of the day, you can know what's going on inside your body. It's your obligation to know. So check out Fountain Life. You can go to fountainlife.com peter to get access to the latest technology to help you detect cancer at the very beginning, at stage one, when it is curable, before it gets to stage three or stage four, and you're a world of hurt. So, Saleem, you sent me an article, a chart. I'm going to just put this up here right now this is our constant debate and we're seeing this again across data center wars in the United States. Data centers are sucking up electricity, driving up the cost for consumers, and also water. Right. It's one of the loudest criticisms of AI right now is that data centers are guzzling drinking water to cool their servers. So this week, this particular chart I'm showing made the rounds and it pairs two figures on one side. Every data center in the entire U.S. according to Lawrence Berkeley National Labs, is consuming 17 billion gallons of water on site. But what it shows is American golf courses that have soaked up 531 billion gallons of irrigation since 2024. That's 31 times as much. And so you know, the posters I'm gonna start seeing on the sides of the, of the highways is, forget data centers, we must ban golf courses immediately.
C
Yeah, Peter, where's the Chinese influence campaign to get America to shut down its golf courses?
A
Yeah, I tell you, I don't see it any place. But here's, here's the, here's the shocking piece of data. Besides golf courses, California almond Farming alone consumes 1 trillion gallons of water 60 times all the data centers combined.
D
So I have one other stat, which is Amazon warehouses occupy 10 times more land in the US than all the data centers combined.
A
Yeah.
D
So it's like such a dot in the buck, a drop in the bucket compared to everything else in terms of land usage, water usage. The human cry is such a completely non data driven garbage bullshit. It's unreal.
E
Well, exactly, that's the concern because the water use is such a non issue. I mean, it's such a joke. But if we take that head on and say, guys, don't worry about water, you know, that the angry crowd is, is going to move to something else equally irrational. So the underlying problem doesn't go away, which is the next issue is going to be something semi sane. This is completely insane, but something semi sane, but still wrong. And then that's going to create a populist movement and the word moratorium. Let's just stop. What kind of a decision, what kind of governance is, let's just stop. But if you look at the history of nuclear and a whole bunch of other things, that's the actual outcome we get. And so David says, I mean, this
A
is the pandemic of fear that I keep on speaking about, that I'm very concerned about. There's an underlying sense that AI and robotics are going to combat humanity, are going to be our foes. And again, I'll just go Back to it. I blame to some degree Hollywood of all the dystopian movies out there. And if all you see is negative visions of the future, you're going to want to shut it down. And what do you want to shut down? How can you shut down AI? Well, you can shut down the data center in your state.
B
Yeah.
C
Also that elephant in this particular room, the Dyson swarm. If the compute all moves to sun synchronous orbit, you can do closed loop liquids including water and other coolants there. But it's not like it's going to be consuming on margin additional water. And then to Dave's point, the complaints, which may or may not be in part the result of an influence operation from a foreign state actor will move to something else. It'll be very low Earth orbit. SpaceX star mines and other competing Dyson swarms are polluting the atmosphere with their. Their decay. Or something else. The complaint will move on to something else.
E
Did you hear the rant about the starship rocket launches earlier? It was Falcon, actually. The pollution from the Falcon launches. Elon was just like, oh my God, I'm going to vomit right now. It was like 0.0001% of all emissions come from any form of rocket launch. But you have to actually answer these questions. He's driving him nuts.
C
That's right.
A
I hope those individuals who are complaining have thrown away their smartphones, don't use GPS, and are just basically going back to subsistence farming.
C
Yeah, as Elon likes to say, let them shake their fists at the sky.
B
I have a fun start. I was doing some numbers around the water thing. It's about 600 gallons of water per Big Mac. And McDonald's sells 2 billion burgers a year, so it's about twice the number of golf courses, the total amount of water that McDonald's uses.
C
Then I can get behind. Okay, so what you're saying, Ahmad, is Chinese influence ops should also be shutting down America. Big Macs.
B
Well, there you go. It'd be a big. It'd be a stab to the heart of America.
E
Yeah, that's right.
A
Definitely improve the health of America as well. Shall we move to one of our favorite humanoid robots?
E
This is so cool.
A
Yeah. So China, as we've discussed before, has gone all in on humanoid robots. It's a national priority. Companies like Unitree and others are racing to commercialize. You know, last report and Alex, we've talked about this. 150 Humanoid robot companies in China under development. And part of their strategy is spectacle and something you're Trying to bring Alex to America. They've been staging public robot combat events literally MMA style. And we've got a video to show. Let me just go ahead and pull this up here of a recent MMA that went viral on the Internet. And it's a beautiful thing. These are only going to get better.
E
You get to watch the full video. It's just the way the fight ends is epically awesome.
A
One of the robots kicks the other robot's head off. You know, it's, you know, remember Rock em Sock em Robots?
E
Yeah, yeah, yeah.
A
As a game, as kids. And so this goes viral. I mean, a lot going on in the robot world. We just saw Hyundai, all of the workers at Hyundai start to strike because they don't want robots brought in on their assembly line. That was fascinating. Alex, take it from here.
C
A few thoughts on this. Thoughts on many different levels. One is mild horror that if anyone who's seen Steven Spielberg's movie AI, where there's. Without spoiling it too much, I think Steven would call it the Dark Sandwich. At the center of the movie, the. The Flesh Fair, where humanoid robots are tortured and abused for human entertainment. I think utterly horrifying. So at one level, I'm mildly horrified that humanoid robots, no matter the extent to which they're being teleoperated here, are setting an inductive prior or bias for future more autonomous embodied intelligences to be basically trying to kill each other, or at least otherwise abuse, physically abuse each other for human entertainment.
E
Not.
C
I'm concerned about that, but one level deeper. Now imagine that these robots are more autonomous, that they're running algorithms that are on the edge, so they're much more encapsulated. And now imagine that these humanoids are in the Chinese PLA. Infantry.
A
Yeah.
C
Because I think that that's the future that we are almost certain to find ourselves in. The west needs to catch up in humanoids. That's why I've supported Pro rl, which Peter, you were gesturing at, which ran their first humanoid robot mini marathon in America in the Boston Seaport a number of months ago. The west, which you.
A
Which you helped. Which you helped organize.
C
Right, Correct.
A
Yeah.
C
Yeah. So the west needs something like this. Hopefully less violent and more economically productive. I'd love to see people cheering on humanoid robots competing to iron clothing or perform some economically productive task and not just kicking each other's heads off.
A
But you prefer the humans to be doing that in the MMA matches?
C
I'd prefer no one to be doing it. I'm not a fan of mma. I think it's destructive to humans and I worry about the message that we're sending to the future light cone by having robots doing instead of humans I'd rather see people in a cage competing if they must compete at all to do something that's positive, some not negative
E
coating like a cage match coding if anything or just sitting there.
D
Okay, so a couple of thoughts. One is my normal commentary around kickboxing is not the greatest marketing demo for humanoid robots but I will acknowledge something here. This is like unbelievably demanding engineering environment. Right. You've got a stress test. It's stressing balance and impact, resistance and recovery and locomotion and latency all like there's about 20 things that they're doing and it's kind of incredible to watch them navigate that. Of course a four armed robot would beat the two arm so I'll just
B
leave it at that.
E
So hello.
D
There we go.
A
So but this is what this is, you know we're going to see this go to competitive sports. We'll see a version of the World cup with robotics questions when people will watch that or not.
D
Yeah, I, I'll say that the real test is whether a human being can make that penalty shot under pressure at that top win the game. Although like watching England implode the other day was really devastating for me but still it's, it's really, I think the, the people much rather watch people in that environment rather than robots.
E
But I think sports is, sports is going to thrive for many, many decades to come.
A
Formula racing, formula racing pushes the edge and I think when we start to see robotic sports it's pushing the edge. I think the point you just made Salim, is important that we're going to see this happening in a competitive fashion so that the top robots and I can't wait to see figure versus Optimus. I think that will be a fun competition, whatever form it takes.
B
Yeah, I think that these robots are a little bit different though. Like I think probably you'll first see the real steel type patelli operated robots because robots can't actually respond fast enough. If you look at the latency of a VLA model like this is impressive from some pre operated flying kicks. But why aren't they doing Kung Fu? When will robots do Kung fu? That's when you move to things like etch silicon, when you move to teleoperation. And I think that'll be the next stage that comes next year. But I think there's a bigger issue that I have with this. Although I love fighting robots and I can't wait to Gundams and all that. These Robots are engine AI T8 hundreds. They weigh about 70kg and they punch four times harder than Mike Tyson. So they could legitimately kill someone. Us fleshy humans. Robots like that should not be allowed on the streets. And there's no regulation against that. You know, like, again, they could be in the pla, People's Liberation army or whatever, but robots are about to enter our household. I mean, who here has a 1x robot on order? You know, like, come on, it's coming. They will be walking around very soon. And we need to have regulations about safety of what the talks are on, these things, of how they operate and others, because they represent a real threat to individuals, because they are machinery. Then beyond that, you will have the embodiment in others. We need to have the discussion of what that looks like when they are autonomous, because these things are delivering themselves by pushing a button on the door, you know, like ringing your doorbell. And the final thing is Unitree has only made 11,000 robots, humanoids total. We are literally at the very start of this. A few years from now, it'll be 11 million a year from 11,000. So we've got to have this discussion fast as well. Lots of talking to do.
A
Yeah, I mean, this is what the work you and I were doing, you know, in terms of how do governments sort of counsel their policymaking around these areas. And it's happening at a blinding speed.
B
Crazy.
A
Yeah. All right, I'm going to move us to the most important conversation we always have, which is the Dyson swarm. And let's take a look at a video from. From our friend Sam Altman.
C
I honestly think the idea with the current landscape of putting data centers in space is ridiculous. It will make sense someday.
B
But if you just do like the
C
very rough math of launch costs relative to the cost of power we can do on Earth. To say nothing of how you're going to fix a broken GPU in space. And they do break a lot still, unfortunately, we are not there yet. There will come a time. Space is great for a lot of things. Orbital data centers are not something that's going to matter at scale this decade.
A
All right, we have the continuing MMA battle between Elon and Sam. Yeah. So fascinating. I'm curious of reactions here, Alex. I'll go to you first.
C
Yeah, I think there's an obvious conflict of interest. We saw similar messaging from Masa Son regarding lack of purported promise for orbital data centers. Remember, OpenAI has retreated from its own data centers. Remember Project Stargate, Stargate project. Stargate has been rebranded from OpenAI owning and operating its own data centers to just leasing terrestrial data center capacity from others. OpenAI is delaying its own IPO. So just not even at the object level. One has to look at OpenAI's messaging here and say perhaps it's not even in a financial or operational position at the moment to lean into orbital data centers, say the way Anthropic, which in their collaboration agreement which was announced with SpaceX AI and for use of Colossus and Colossus 2 far friendlier to Orbital Data center based compute. So I think the crossover is going to happen. Elon's messaging regarding when this crossover is going to happen is two to three years. You see other analyses that suggest that the unit economics for orbital versus terrestrial data center costs are going to cross over sometime by the early2030s. I'm not sure which is the case, but either way I think there is an obvious conflict of interest. And just as we were discussing with Philip Johnston, barring some surprising left turn, I expect that OpenAI's tune is very conveniently going to change on ODCS sometime in the next two to three years. Right on time.
A
And of course Elon's response to this is we'll be launching them in two years, so stay tuned and watch.
E
Well, I think, I think anyone listening to this video would say, okay, Sam says space data centers make no sense. Elon says they make sense. The two guys hate each other. But if you actually listen closely to Sam's words, they don't disagree at all. Sam is saying that space data centers will not be meaningful this decade. There will come a time, but this decade's only three and a half years left. And if you look at Elon's forecast of his launch rate, they agree, actually. So they're just hating on each other all the time. And it seems that way in this phrasing, but the truth is pretty clear. They both have the same numbers. So Alex is right. They're going to space. It's going to take a while. I think a couple percent of all COMPUTE will be in space by the end of the decade because we're building out on land. You if as quickly as we can too.
A
And the ocean.
E
But then the lines cross.
A
Yeah, you know, Alex, you and I were going back and forth texting while the starship attempt. Starship 13 flight was making an attempt a couple of days ago and it's been rescheduled. When this pod comes out, we'll be seeing a next launch Attempt on Starship 13 on Monday of this coming week. That launch was thwarted at T minus 0 when two of first time I've
C
ever seen that, by the way. Zero.
A
Here's the point. Two of the 33 Raptor engines on the booster stage of Starship did not ignite and they're going to be replaced. But here's the extraordinary point. So, by the way, SpaceX's stock dropped 5% on news of that failed launch, which is kind of ridiculous. The point people need to realize is that was an amazing demonstration of technology. The fact that you could shut down at T equals zero safe, the vehicle, unload the methane and the liquid oxygen and that's an. You know, I was part of the space industry in the 90s before it was a space industry. And those vehicles would have exploded on the spot. Right. They would have failed on the spot. The ability we have to control them at that level of detail is evidence of the extraordinary engineering that SpaceX has done.
D
I thought that was the most interesting part, which is how, how quickly the system diagnoses the problem and returns. It would have taken months and months to do this and fix it and recover everything and replan another launch. And you're like, yeah, we have problem, Shut it down, redo it. Oh, we're starting Monday. I mean, it's amazing.
A
Yeah, extraordinary.
B
I think if you're serious about spending intelligence with what we know, you have to have a space play. Open Air is going to buy like Planet Labs or something like that, you know, like, then the tune will change.
A
All right, I'm going to go to some AMA questions. So, Imad, you had suggested I post questions to X and we have a number of questions coming about Kimmy from our X audience. Let me go ahead and show these and let's dive in. So Imad, I'm going to give you first crack. Which of these questions do you want to answer?
B
I think probably number four is an interesting one. Given Kimmy K3's lower token efficiency, is it actually as cost effective as advertised compared with Sol or fable? So Kimmy K3 is an expensive model relative to the other Chinese models. Like Deep seek is now a dollar per million tokens. Kimik 3 is $15, Sonnet is 20, or Opus is $40 and I think like Fable is $60. But that's because they're actually making money. When you back out the numbers from the Chinese models and the chips they're running on, they're probably making 80, 90% margins now. And that's with their Chinese Chips which aren't that efficient. For running this, we will see the cost of K3 drop by 10 to 50 times I think in the next few months as it gets optimized. And right now it uses twice the number of of tokens for the same task versus GPT 5.6. Again, a frontier model that uses 37% less tokens than 5.5 or Fable Again, we're going to see that drop because everyone and their dog is going to optimize the crap out of this like you've seen fireworks. Just raise at a $17 billion valuation. Others like modal at 10 billion. Base 10 at 10 billion. These are the inference providers of open source models. They've all raised a billion dollars that they're now going to spend to optimize the Chinese model and make it more efficient and run it. And so American labs who do the inference side of things are going to optimize the crap out of this. So we will see it catch up.
A
All right. And by the way, I welcome the mates to lean in on these questions. But Saleem, you want to go next?
D
Given that I made the comment about number one, how much could Chemi K3 devalue US Frontier models? I'll stick with my original estimate of about 75%, 50% from the US regulating the front end. And then you've got lack of compute so on the supply side plus the front open source models kind of within a release, barely of where you are. That bleeding edge is such a perishable thing. I would say 75% drop. So if you OpenAI was worth a trillion bucks, I'd put it at 250 billion. You still have a very valuable business because now the competitiveness you have to compete on reliability, security, integrated tools, ease of deployment. But the actual frontier cutting edge, it becomes one ingredient amongst the whole thing.
A
I would not want to be inside these frontier labs right now. It must be a frenetic code red 24 7.
C
It is a total rat race. I have so many friends at the Frontier labs. Friends who are jumping hypothetically from one frontier lab, Google, which is nowhere at this point, missing in action to other frontier labs. It is a total rat race.
A
Yeah, it's crazy.
E
Dave, you have a choice for me?
A
No. Pick one. You got two and three, I think.
E
Okay, I'll take two. What does the release of Kimik 3 due to the open source versus closed source race, will this force the large companies to provide more product? I think they're implying more open source product. Yeah, it's a total Game changer in the sense that anyone with resources can build an internal model that's tailored to a specific use case and then use it as a defensive moat. I don't think the large US model providers will go open source. I think they're committed to their pathway. So if you were talking to Anthropic right now, they would say, look, Kimi has caught up for a week, but Fable 5.1 is coming out in just a few weeks. When you look at the all important enterprise use cases. So white collar automation, drug discovery, people are going to use the best model no matter what. If you're using an AI to design a car or a rocket, a slight improvement in the design has massive payoff. So you're going to use the best of the best of the best model. So the Anthropic guys are going to scramble to stay a step ahead and keep their price point nice and high. The cost of the model itself is so small compared to the benefit that people will pay the price. So it does create. Like Alex was saying, the rat race is incredible. But people aren't going to switch to Kimi unless it's proprietary data they want to keep in house and they want to tune their own. Or Kimi actually bypasses Anthropic, which it hasn't done. You know, it's only caught up, or not even quite caught up.
A
All right, Alex, number three.
C
All right, number three asks, and these are, I think these questions seem to all be variations on a theme, but it asks, how can US models. I think this means US Frontier model providers continue to justify their massive valuations if China can leapfrog with an open weight model at less than half the token cost. So I don't think the premise is quite accurate. There are so many elements, so many layers to superintelligence, and quite frankly, superintelligence itself is as it fully develops, I think, far larger than the total GDP of the entire world. Anyway, there is an enormous amount of pie that can be sliced. But to the extent we're talking about, say, Google, which as I was mentioning earlier, seems to be MIA at this point on the frontier. I can't find a single top Google model at this point on the cost frontier for capabilities, what does Google do? Well, they can continue to race, obviously in terms of capabilities, but if I'm Google, I'm thinking, yeah, I want to become a hyperscaler. I mean, Google obviously is a hyperscaler, but a hyperscaler provider to other frontier labs, that's one obvious venue of differentiation. And we've seen, we've seen that approach vector from SpaceX AI itself which has now signed deals with Anthropic. We're seeing it with Meta, interestingly, which on the one hand is offering Spark 1.1 and on the other hand in the past two days, just as we were going to air, it was announced that Meta is exploring selling $10 billion of compute to Anthropic. So differentiating by going downstack and offering your compute up to other more competitive providers, whether Western, usually Anthropic, sometimes OpenAI or Chinese models in a self hosting model that's one area you can also go upstack. You can try to vertically integrate and offer applications that are being that are benefiting from the commoditization of their complement, namely the model layer. You can also I think the premise that valuations somehow are going to net shrink just because Kimik 3 exists now is completely fallacious. We saw that incorrect thinking happen with the original Deep Seek Shock which was at the time also branded as a Sputnik moment. So we saw a bit of a hiccup in capital markets at the time. But as always Jevons Paradox kicks in and we see the value of chip stocks ultimately increase, not deflate and we also see it's open weight. So there's absolutely nothing in kimik3 that OpenAI and anthropic and other Western Frontier labs can't just immediately reappropriate for their own internal models.
A
You don't think that the amount of revenue these labs are going to make because gets reduced as people start to use Kimi K3 for their work instead of the API calls?
D
No.
C
For example, so I spend and my portfolio companies spend an extraordinary amount on let's say Anthropic and OpenAI and to my knowledge, my expectation is Moonshot would have to release a 2x3x10x better model than say Fable 5 to have a massive diversion of that spend right now what what K3 buys at the moment, to the extent it's legal, query how much longer K3 will be legal to host within the US but assuming it remains legal and regulatorily uninhibited, all it results is greater in house self hosting, but it's not at the top of the frontier. To Dave's earlier point, Fable 5 at the moment is so if you're trying to do like solve the frontier of problems, K3 is not causing you to divert your spend.
A
Well let me hit that point. You just made Alex and ask you and the other mates a question here which is, do you think it's possible that K3 that some legal policy in the United States prevents US companies from downloading K3? It's gonna be on the open Internet, it's gonna be available through a multitude of sources. Beyond Hugging Face, can it be shut down in the US it can effectively be shut.
C
This is not prescriptive and I'm not a fan of this policy but I think it can effectively be shut by requiring that every public corporation disclose any use of Chinese open weight models and subjecting them to scrutiny. As we were going to air the latest we talked in the last pod about Demis proposal to create a FINRA like entity that would regulate the frontier. Well guess what the reports are that the present administration is actually running with a proposal like that and is planning to or at least expect exploring creating a FINRA like agency to regulate frontier AI that would live under the sec because the SEC already has statutory authority to operate FINRA like industry advised and funded entities. So it's a natural place.
A
Self regulated organizations.
C
Yeah, yeah. Self regulated governance AKA regulatory capture cartels under the sec. And so I think it's completely plausible, albeit I think highly undesirable that we get sometime in the future an SEC sub org that looks like finra. That basically makes it completely economically infeasible for corporations of any size, especially public corporations, to actively use Chinese open weight models.
A
Any other comments on this?
D
I've got comment on this. I mean this is ridiculous in terms of trying to limit the the use here because once you release the weights, right, stopping them, you can mirror them across jurisdiction. You can use peer peard networks, hello VPNs. All you're going to do is deny American researchers and startups access to those models and security experts while the rest of the world goes ahead on building on those models. I don't think there's a viable approach.
B
I mean this is the same.
A
Yeah please.
B
This is the same as denying Americans cheap insulin. I mean it's again regulatory capture, right? Like why can't you have generics? Because again you have the regulatory capture point. There's operation. I think they're calling it Gold Eagle. To approve access to frontier models you will have anti token laundering regulations. You'll have know your prompter regulations. Like the US Government's really realized that this technology is about to break through. And I think that they're a lot more worried about it than China is. You know like China again you look at that Xi Jinping speech. I would urge everyone to kind of check it out they're like full on open source. We're going to do this. America doesn't know what it's going to do. But as you said, there's a real chance that they might hobble American capitalism. And oddly, China's encouraging capitalism. It's going to get very.
C
The CCP saves American capitalism from itself. It's a crazy future.
D
The world is so weird.
A
All right, let's go back to you, Salim, on next question.
D
Okay.
A
Which some of these are a little bit duplicative?
D
Yeah, yeah, I'll take number five. Would you trust Kimmy K3 to write your code for you without oversight or review? The answer is no. But I wouldn't trust a human being to put consequential untested code into production either. Right. The question is not whether we trust the modelers, whether we trust the development system around it. So, you know, AI generated code needs to be run in a sandbox and pass automated test and security scanning and all sorts of things before it goes into production. And then you do proportionate permissions based on the use case and on the potential impact. This is the same thing we talk about. Whatever the workflow is that AI is running, you're still going to need human review at the highest level and at the highest consequential inputs. A lot of the routine can be automated. But the scalable model is not AI with no oversight. It's machine generated plus verification plus human accountability combined. That's going to give you the real power.
A
All right, imod?
B
Yeah. I think what role, if any, distillation play in K3 development? They distilled data clearly from Opus and others, but to be honest, using Kimik 2.5 and Kimik 3 now quite intensely, it feels different. So I think they did a lot of their own data creation based in part from distillation, but everyone's distilling from each other right now. The one area that it's clear that they've had a big leap ahead is in the front end development. Again, this isn't the best mathematician in the world, although it's quite a good general model. It's not the best cyber attacker from our benchmarks, but they've kind of done something original and new on the front end game, consumer entertainment side of things, which I think is really interesting, although that might be also because it's a multimodal model.
E
Dave, number seven, what are the reasons why Kimik3 might not be as good as advertised or we shouldn't use it? The scenario where it's not as Good as advertised is if it's benchmaxed and you know, in two weeks, you know, the open source will be out, we'll have beaten it to death. We'll know the answer if they benchmaxed. So we're going to find out. I think it's unlikely that it's benchmaxed to the point where every company in America right now should be in the world right now should be saying we need a crash program with our best possible advisors to decide are we going to do our own model on our own on PREM hardware or are we going to use anthropic or OpenAI or Google and just trust that API. But we need to decide whether tuning and training on our own proprietary data gives us a long term competitive advantage. And so there's going to be a desperate shortage of good advice on this and vendors and McKinsey consultants and you got to grab those resources, Exo consultants, you know, make seed stage investments, get your network together, find out who can answer that question for you internally on your business and your use case quickly and then commit to the path. And you know you can, you can do something internally and still use the APIs, but if you don't start down the path of evaluating Kimik 3 on your own, you can't really come back to it later. So I think everybody's got to just get going on this question. We'll know in a couple of weeks though, whether it was benchmarks to hell or not. But I think it's very, very likely that the open source path is a viable path for every US and world company and government.
D
Now can I just add to that real quick?
A
Yes, of course.
D
Very simple suggestion. For every company, implement two installations, Kimik 3 and Inkling, Fine tune your own internal data, because that learning loop is going to be the proprietary goal that you don't want to lose. And start there.
A
Alex, why don't you close us out here? You've sort of answered number six already, but perhaps you could expand on it.
C
Yeah, I'll say something new. So question six asks, should the US move to block loading the weights of the next Kimmy release onto hugging face, I'll give a conditional answer. I think that if some party presumably in the US can prove to a cognizant court that the next Kimi release, presumably a reference to this Kimi release, was somehow obtained or derived illegally, maybe through copyright infringement or illegal distillation of traces or something like that, that would probably be grounds for blocking its release in the US but, but if no one can prove that Kimmy's parent moonshot did anything otherwise wrong in creating it. No, I don't think the US should be blocking its release in the process. I think if anything, quite the opposite. I think every US Frontier lab should be closely scrutinizing it and learning whatever they can so that we can leapfrog it. And I would like to see far more outward pressure from US labs creating the best in world open weight and open source models so that it's not the ccp with their new belt and Road for AI initiative blanketing the world. Some would even say dumping superintelligence on the rest of the world or the so called global South. It should be the U.S. the cannon of freedom, the arsenal of freedom that's also the arsenal of superintelligence. Showering the rest of the world with open weight and open source superintelligence, Not
A
China showering the rest of the world. I love that. And remember, we're moving towards intelligence too cheap to meter, but a million times more available and more powerful than ever before. Everybody listening. I'm grateful on behalf of the moonshot mates here for your time. If you haven't subscribed, please do. We're gonna be putting this out more and more often as we're starting to see the release of states move from months and weeks to days. And there is no time to sleep during the singularity. Gentlemen, what's in store for the week ahead? Imod, I'll go to you next. Yeah, imod, what's. What's news in your life?
B
Yeah, just getting a whole bunch of research papers ready to release. So finally it's going to be exciting.
A
Yeah. Again, more acceleration for intelligent Internet. Your company.
B
Yes, yes.
A
Incredible. Celine, please.
D
Tuesday I have my next meeting of Life session, 7pm Eastern. For those that are interested.
A
Where do they go to find out?
D
We'll have the link below, but it's openexo/openexo.com mol so if you, if you've
A
not participated in one of one of Saleem's Meaning of life sessions, they are extraordinary. We'll take you beyond the AI into the realm of philosophy in theology. Alex, are you traveling? Coming up. What's going on with you?
C
I'm so focused at this point on literally solving everything. I'll say large swaths of the sciences at this point I'm convinced are so thoroughly cooked. More to come on that subject. Peter, you and I wrote solve everything about it, but now it's actually coming true.
A
Yeah, I'm excited you're going to be doing an AMA with my abundance community coming up. That's going to be a fun deep dive. And of course we're going to have you during the moonshots gathering on September 25th. Doing in. In fact, all of us will be here. Imad, you're joining us in LA in September.
B
Yeah.
A
It's going to be fun to have all of us together again for the full day. Dave, you know, this has got to be the most exciting time to be in Lingq Studios.
E
Oh, my God.
B
Yeah.
E
I think that discussion we had of quantization on this podcast that Imad kicked off, I think that now vaulted to my new best piece of media ever recorded passing Leopold Aschenbrenner. I got to go back and listen to that again in slow mo. And also, you know, we had Vlad Bulovich from MIT Nano in this week. He's going to advise and help us on our new. Our new startup working on photonic computing. And he gave us a whole roadmap of people I need to meet next week. So we are looking to add two MIT people with our Princeton team to work on just the photonics, quantized photonics side of the equation. So I'll be working on that next week, but I think I can take that video we shot earlier and use it as a recruiting tool. It was just so freaking brilliant. You guys are incredible.
A
Yeah. I love you guys so much. What a great week.
D
Awesome conversation.
A
We'll see what breaks tomorrow. Yeah. Over the weekend. Emergency pods.
C
We need emergency pods every day by January.
A
All right, be well, everybody. Thank you for tuning in to Moonshot, your front row seat to the Singularity. Take care, guys.
D
Peter, awesome job as always.
C
Thanks, Peter.
F
This episode is brought to you by Accenture. When your advertising operations fall out of sync, everything else follows. Spotify and Accenture are working together to reinvent the rhythm of ad sales, using automation, analytics and smarter workflows to simplify campaign delivery and access better data across the business. The result? Less time spent on operations, more time connecting brands with the moments and fandoms that matter most. Learn more@accenture.com Spotify.
Urgent Update: AI Sputnik Moment – Kimi K3 Released ft. Emad Mostaque
Date: July 19, 2026
This high-stakes emergency pod convened in the wake of the release of Kimi K3, a breakthrough model by China's Moonshot AI, which represents “America’s AI Sputnik moment.” Host Peter Diamandis and his full Moonshots quintet—Alex Weiser, Dave Blunden, Salim Ismail, and Emad Mostaque—convene to dissect how Kimi K3 shattered expectations, leapt to the top of global AI benchmarks, and is triggering a new phase of global AI competition. The conversation unpacks the technical and geopolitical implications, forecasting a world where intelligence is abundant, open, and accelerating beyond regulatory control.
Speakers: Peter Diamandis (A), Alex Weiser (C), Dave Blunden (E), Salim Ismail (D), Emad Mostaque (B)
00:41–02:17
Key Segment: 04:26–10:38
What happened:
Domains topped:
Impending Open-Weight Release:
Notable Quote:
“They didn’t just close the gap, they jumped the fence overnight... Frontier Intelligence is now a totally perishable asset.” – Peter, 04:26–04:41
Key Segment: 05:57–12:02, 32:21–34:31, 62:08–72:32
Technical Highlight:
“The fact that a recognizable transformer can almost match GPT 5.5 max on the cost-performance frontier… is pretty striking.” – Alex, 07:28
Data Curation & Efficiency:
Cost Parity & Open-source Shift:
Key Segment: 14:23–23:39, 26:12–29:43
AI Race is “Unstoppable”
Impact on U.S. and Chinese Policy:
Enterprise & National Security Angle:
Notable Quote:
“Any government or enterprise interested in the very latest cutting frontier model doesn’t have time to evaluate it. ... Now all the value comes in architecture that can swap models.” – Salim, 17:55
Key Segment: 21:47–31:54, 109:10–116:16
Revenue Model Disruption:
Valuation Observations:
Key Segment: 62:08–72:32
Notable Moment:
“This allows every device in the world to have built-in persistent intelligence... this decentralizes capability at the most infinite level.” – Salim, 69:46
Key Segments: 73:32–83:33, 55:31–59:44
AI Superforecasting:
Management, Wisdom, and Human Value:
Entrepreneurial Call to Arms:
Key Segment: 115:05–118:14, 29:08–29:43
Regulatory Capture Risk:
China’s Open-Source as Global Soft Power:
Key Segment: 93:15–100:51
Key Segment: 101:08–106:10
Key Segment: 88:00–93:07
Panel’s Final Take:
“There is no time to sleep during the singularity. Go build. Find your massive transformative purpose and make a dent in the universe.” – Peter, 124:03 onwards
End of Summary
For detailed breakdowns and explanations, reference key segments and timestamps above. Panel recommends listening to the full episode for deeper dives on technical themes, economic forecasts, and actionable insights.