
Loading summary
Host
We're here with Akshat of Modo, CTO of oto, together with vibu. Congrats on your cvc.
Akshat
Thank you.
Host
Your party yesterday was amazing.
Akshat
Yeah.
Host
All the photos and all the swag.
Akshat
We had a bunch of art installations, which is kind of fun, seeing our products on pedestals next to Rodin.
Host
Very nice. Very nice. When you started, it was not the GPU Inference company. I mean, maybe it was in your mind. Take us back to the origin story.
Akshat
I actually first met Eric, who's the CEO, through an investor, and back then Eric was already thinking about building a new kind of runtime. And he got there thinking through, why are workflow orchestration products so hard to use? It's because you have to run them on Kubernetes. Kubernetes is hard to manage. It's not built for burstiness and custom images and the terrible developer experience.
Host
And I'll interject for listeners who are new, we interviewed Eric two years ago and there's a bit more of the story there from Spotify and all those things. And I actually came across Eric through a data console because he did that talk on the serverless container stack that you guys did, which is like, that was my first, like, okay, I need to take models very seriously moment. But it was still very unclear. Do I actually need all this for just my data pipelines?
Vibu
Yeah.
Akshat
I mean, initially what we were thinking about was if we build a better runtime. It's a very useful primitive in itself. There's a lot of things that get solved by serverless functions, like you can do ETL stuff, you can do job queues, you can do all this bursty processing, which it turns out every company had, needs more. But then we also were thinking about this as like, this is a primitive that we can build a whole collection of products on, which are very verticalized. So perhaps data engineering would have been the first one. But we were thinking about inference back then. It was more classical inference, like computer vision stuff and running XGBOOST and whatnot. But we added GPUs to the product a year before ChatGPT came out.
Host
Nice.
Akshat
We just didn't think it would be that big of a deal.
Host
Yeah, just like add a 100.
Vibu
Was there any early key problem that really sparked off why you built it?
Akshat
Yeah, primarily. It's just none of the tooling that was out there was built for one, a really great developer experience. And also there's a general trend of. A lot of the workloads that we were seeing were very. I wish There's a better word for it but compute heavy like they need one need a lot more resources. You need to burst up and down a lot versus kubernetes designed for slow scaling and more for web server use cases. And also there's just a lot more specialization in what kinds of environments these workloads run in. Sometimes they need accelerators, sometimes they need different kinds of images. And this is just a consistent thing that we saw across a lot of companies. That would be the next step.
Host
Yeah. Be nice. I don't know how much this factored into the early story, but I wrote a post when I was at Temporal about infrastructure, software defined infrastructure or something like that.
Akshat
The self provisioning work.
Host
Yeah, I can't even remember my own post. And then you put me on the. The landing page.
Akshat
Yeah, we really like the term and so we stole it.
Host
Because you had the insight that everything can just be in decorators next co located with the code. Right. Was that a big part of the original story or it's just like a
Akshat
DX layer that was really important because we really didn't want people to spend so much time writing YAML. And it seemed like you could really condense the surface area of what you're doing, put in code so you can actually operate on it just like you can operate on other code and build stuff that's more expressive and dynamic. And so yeah, that was always a very important part.
Host
Then the pushback is this is a dsl, it's your closed source. I am locked in to modal.
Akshat
Yeah, we never really got pushback for that because the nice thing about modal is you can bring whatever code you have. And sure the DSL is at the sort of configuration there for what hardware you're using, how you're scaling things up, but you still own the code. And that's been an important part of our story even as we do inference now.
Host
Yeah.
Vibu
How much of it do you think still stays the same today? Like if you were to build something today, devex obviously very important, but I feel like a lot of this has kind of been changed with just hook it up to an agent, have cloud code, have CloudX implement the tool. There's very agent native primitives that are kind of different than if I'm doing this myself. Right.
Akshat
We've actually changed our SDK team to think about agent experience, developer experience. And we think that the same benefits that apply for DX also actually apply for ax, which is why would you have an agent read through hundreds of kubernetes files and like write YAML that's not even typed when it can basically make a couple of changes in a decorator and it gets this sort of self provisioning runtime of being able to see its changes live in action. Yeah. It just seems from the customers we talk to, they actually find the model is way faster to use for agents versus operating on a different substrate.
Host
Yeah. Because again you co locate the infrastructure requirements to the code that runs it. Well, the negative thesis now is that nobody's looking at their code anymore, so there's no point.
Akshat
Yeah, I mean people are looking at code. One thing we actually still see is really important is observability. How good is your dashboard? And of course we push a lot of it to the CLI so the agents can do their own investigation. But you still need humans to go interpret what's going on and make judgment calls and whatnot. And that's I feel like maybe more important now than looking at the code itself.
Host
Yes. Because you can try to treat the code as a black box and then use sort of see the observable action that comes out of it and then just prompt to change. So I think actually it takes a bit of restraint to not specialize to say I want to ship a new primitive and then just be general purpose. People ask you what are you for? You're like, I don't know, we can do this, we can do that.
Vibu
I'd be curious to say, okay, if we were to ask you what is model for? Even at a high level there's a lot you guys do sandboxes, CPUs, everything. How do you answer?
Akshat
Modal is a cloud platform that's built for where we've built the primitives from scratch for AI applications. And right now it basically covers inference, training, batch processing and sandbox workloads. But we're building a lot more.
Host
And I noticed you didn't say web server. So there is still a role for the always on large scale Kubernetes type things.
Akshat
Yeah, absolutely. We're not trying to compete with the renders of the world because we think the differentiator for us is our other. Workloads that need specialized compute need to scale up and down a lot. They're shaped differently.
Vibu
I think you're building a lot of it alongside the startups. They're innovating quite a bit. Even in your latest blog posts, even in the series C, the customers that you mentioned here, the cognitions, decagons, ramps and whatnot, they're innovating with you. Right. And that's not something AWS is doing directly with.
Akshat
Yeah, absolutely. I think. I mean, this is again, classic. We're small team, we can move really fast. Our engineers are working with our customers and figuring out.
Host
So my first week at cognition, I walked in, there was someone wearing a moto shirt. I was like, what are you doing here? They're like, yeah, I just. I am embedded inside of cog.
Akshat
Yeah, I think that was Peyton. We sent him over. The latency of communication is too high otherwise.
Host
Yeah, Distributed node. You have to place one and co locate, actually. So I had a direct personal experience. So I worked on small developer three years ago. It was inspired by Claude 1. I think you onboarded me at some point, just before and I was like, oh, I need some bursty compute. I was just going to try using modal and it was a pretty pleasant experience. Apparently I showed up in the board meeting. I like the analytics.
Vibu
Yeah.
Akshat
You blew up on Hacker News and we got a big traffic spike. I think the way you use Small Developer was modal functions for running stuff, which was like. That was a good use case.
Host
Yeah. So to me, that was proto cognition.
Akshat
Right.
Host
If only I had stuck to it. That was like, if you just draw the tech tree out, you're like, yeah, probably this will happen.
Akshat
Yeah. He was so close.
Host
I didn't realize.
Akshat
But the funny story there is. At the same time we were talking to a bunch of customers who needed something like sandboxing. This is like 2023.
Host
Yeah. So we built a new API right after that.
Vibu
Yeah.
Host
Yes.
Akshat
Like we built sandboxes in May of 2023 before anyone was even knew this was going to be a thing. And the first example we published was we took Small Developer and put it in a loop so the agent can iterate on itself.
Host
Loops are hard these days.
Vibu
Loops in. When was this? 2023.
Akshat
A small check. Mid 2023.
Host
Obviously for listeners. The problem was the models are not built for any of this. They're not post trained to understand looping and cell correction and tool calling was there, but also not that great.
Akshat
Yeah. I don't remember if you use tool calling in this one, but yeah, the models would just diverge after 10 iterations and not produce anything meaningful.
Host
Yeah, but like then. So I mean, okay, like now talking to myself three years ago. The answer is collect all the failures, build benchmark and then collect all the, you know, examples, build the RL environment, sell it for like $10 billion to Meta and then. And then also train a model and then sell that for $60 billion. To Elon. And this is for that money machine. It's actually about that hard.
Akshat
I mean, it's hard to have that kind of inherent conviction that the stuff will get that much better.
Host
In retrospect, it's so fucking obvious.
Akshat
Fair enough.
Host
What else were we doing back then? I don't know. Anyway, yeah, that was the start of your sandboxing journey, right? I feel like it didn't blow up, blow up until last year. Yeah. So there was like a couple of years of quietness.
Akshat
Exactly.
Vibu
We were very underrated. Product value, like my experience with Modal. Charles, before he had joined Modal, met this guy at a hackathon and he really insisted we wanted to run some small model not hosted anywhere. And he's like, ah, there's this cool company, Moto. They'll like spin up a GPU sandbox. We can throw it on there. They'll take a hugging face link and there's so much value just right there.
Host
Right.
Vibu
Like instant hosting. Spin it up, spin it down. It'll stay cold. But, you know, we run the demo a few days later, it'll come back up. And like, all this stuff in retrospect is still what we needed, like today.
Akshat
Yeah, I mean, it's still needed today. Obviously, workload shapes have changed a lot as we run stuff for people with really massive production scale. And there it's not about scaling from 0 to 1, but it's how do we scale really elastically from like 1,000 to 1,500 GPUs very quickly in a given region? It's the same shape problem.
Vibu
Okay, so you look at, say, cursor composer, right? We'll do RL on a model every couple hours. You guys have a whole version of rl, inference gym and whatnot. When you look at workloads like that, you're basically doing train runs where you need to scale up, scale down every hour. Thousands of GPUs. Right? That's the example for we do need it.
Akshat
Right. Well, actually, I'll take a step back and maybe talk about how people use model today. Because our biggest use case actually is elastic inference. And the thing we first found product market fit with was inference for custom models. So we kind of stayed away from the LLM space. And we were serving companies like Suno for audio, Runway for video, robotics, Comp Bio, companies that train their own model elsewhere. But Modal is the best black box for deployment, scaling to however many GPUs you need as your traffic pattern changes. And we saw all of them actually have a Very unpredictable traffic pattern. It's like diurnal, it's. And some days the company will launch and they'll need way more. And it's not just one model that they deploy. All these companies deploy lots of different models in different regions. And so the auto scaling problem becomes even harder because then you have to scale within a certain region and those cycles sort of are offset. So different times you need to scale up in different regions. So that's like our sort of.
Vibu
And that in and of itself is a huge category. There's a bunch of inference providers which provide this. Fireworks does this as a service together. Whatnot. Base 10 that's kind of carved into its own niche for language models, at least right now.
Akshat
Yeah, I mean the thing that we have actually specialized in is the auto scaling aspect because we found that it's not universally true that everyone else can auto scale. And we've gone deeper into it on the tech side by side. We've incorporated GPU snapshotting into the product so we can actually take the GPU state like your Torch compiler model, snapshot it and next cold starts way faster. And so going back to your question, that's why you need a lot of burstiness for inference, but then people also do a lot of on demand training. Like for RL stuff, your rollouts are bursting. As you said, people also do a lot of batch jobs. So we'll see a lot of companies before they have a training run. They'll need thousands of GPUs around encoding or something like that. And I think those things are much more bursty than I agree that agents are not that bursty sandboxes are. Except when you're doing rl, RL is insanely bursty.
Host
Yeah.
Akshat
Yeah. Like when you're doing rollouts, you sometimes need 100,000 sandboxes.
Host
Yeah.
Vibu
I'm curious if you've seen early sparks of continual learning. There's some people, like our friends Engram recently announced this. They're trying to do training that also seems like a different workload. Right. If you're doing training 24.7per se, there's a very weird dynamic of how you're using GPUs between people and whatnot. But seems like something you guys would work for.
Akshat
As you said, we're fortunate to work with a number of customers at the Frontier and grab some of our customers and they are taking the primitives we have and trying to use them in very interesting ways. Like continual learning. It's possible as the stuff gets better, some of that will be Part of our offering as well, if more people need it, but we're just waiting to see how it shakes up.
Host
Is there a primitive that you added after sandboxing? That was the next step in the story, I guess.
Akshat
We've been going much deeper into LLM inference because we realized that some of the advantages we have with auto scaling, again, especially in different regions and whatnot, are not present elsewhere. And the place where we had a gap was we weren't working on the model layer itself, we were a black box and we realized that we actually can get to frontier level model performance by having great people who work on all this. And we've actually been open sourcing a lot of our work in terms of recently we shared our work on D Flash, which is a block based speculator and we've open sourced all of it. So you can get by using open source D Flash you can get the same performance as you would with one of the proprietary providers. And the next thing we're thinking about here.
Vibu
I thought this was actually an interesting blog post as well. Right. I think in here you make a claim that. Not a claim, just that how effective speculative decoding really just get to anything you want to point out from this around what people should know?
Akshat
Yeah, absolutely. I mean the high level summaries would have helped to describe what speculative decoding is.
Host
Yes, I think so. We've covered ego and all this hydra and all those things, but it was like two years ago.
Vibu
I think it doesn't hurt. Right.
Akshat
The speculative decoding is you have a smaller model called a draft model, predict tokens ahead of the bigger model and then you have the bigger model verify all of this, all the tokens are predicted. And the reason it's faster is if you're predicting one token at once, you're kind of bound by memory bandwidth. But if you can batch the verification of the draft model, then you're much more efficient using compute and it's faster. And as long as your draft model is producing a lot of tokens that can get accepted, which is called the accept length, you can get speed up. That's multiple times of the original model speed. And that's what we highlight here. People talk a lot about we made these kernels faster and whatnot. But improving kernel only give you a few percentage points of improvement. And increasing accept length literally is multiplicative decrease in 2 to 4x. Yeah, exactly.
Vibu
Without much head on performance.
Host
Yeah, I think it may be. I mean you are running a second model, right. So maybe it's more expensive in the compute. But I meant quality performance.
Vibu
But yeah, I mean, I think so
Akshat
there's no drop in quality performance because you're always, you're never accepting a token
Host
that's better or same.
Akshat
Exactly.
Host
Right, yeah.
Akshat
And so we've been working a bunch on dflash, which is a block based speculator. So instead of predicting one token at a time, it's predicting a block and we've been open sourcing our work with it. The next thing for us here is for helping people train speculators and custom models. It's something that traditionally is very FDE driven, support, deployed, engineer driven, make you work with customers and help them do that. And our vision for this is why we launched auto endpoints is we want to make frontier level performance available to everyone. And so we mentioned the announcement, we kind of teased it. The next thing we're launching is basically as you run an auto endpoint, we shadow traffic and do you want me
Vibu
explain what auto endpoints are? Yeah, lovely.
Akshat
Yeah. So this is I guess going back to your modal is you touch the code, but sometimes people actually don't want to touch the code and they want to get started with an endpoint that works and has all the great performance and scalability that modal has. So we've made that easier with basically a way to create an endpoint from our ui, from the CLI that has all of our optimizations that we talked about, like the D Flash stuff already baked in and there's full transparency. So we give you the code, you can go run it yourself and if you want you can eject out into the full modal experience, which we see as people get sophisticated, they do want to tweak the models, they want to fine tune stuff. You can still do all of that, it's not a black box. And yeah, the next thing as we tease later in the post, is how do we give you value even beyond this in terms of having your draft models evolve as your data distribution evolves again without having to talk to a person,
Vibu
I guess just to kind of understand it directly. I mean, you know, obviously you have the GPUs, you have an endpoint that's compatible, you serve open model. If someone was to do this themselves, what's the delta that you guys provide? So you do a lot of open source, great work on effective inference. How does it compare to say I take the same model, GLM 5.2 FB8 take off the shelf inference engine, VLM SGLang, get compute of similar capacity, similar cost, what's the Kind of delta that plugging into something like this offers. Outside of the benefit of scaling, it's
Akshat
kind of interesting because we've taken the approach of open sourcing our contributions and upstreaming them. We work closely with the sglang team. We actually want the improvements that our team comes up with to be there and open source for others to use. Even outside of modal, the benefit to us is we have a team that has significant expertise in terms of if you do have something that is not there, our team can help you get that performance first. The other thing is with these endpoints, we are way more elastic, as you said, than anyone else. And you have true scaling to zero, you have true burstiness and in practice that matters a lot more to people than just finding the GPU and running model code.
Vibu
Yeah, and I will say it's actually not that straightforward. What I said is easier said than done.
Host
Right? Yeah.
Vibu
I think still for the average person, still hard to just gut check using different. There's quite a bit of combinations you can make there. The trade offs aren't really known at face value.
Akshat
Yeah, I mean it's not just that. I think it's that running production grade inference is a hard infra problem. Even if you subtract out the auto scaling is controlling things like tail latency and making sure every request is delivered at least once and whatnot.
Host
There's a lot of innovation that you can do here. I think it's very interesting that you're starting to encroach on as you become a full cloud, you're starting to encroach on other people's turf. What will you not do?
Akshat
Well, we want to follow our users and make sure they get a platform that has everything that works well together. So right now we're kind of focused on the model lifecycle and the agent lifecycle. So both like going from data prep to training to inference. And then also if I want to deploy a background agent, let's say, you know, sandbox, super storage, a whole bunch of other stuff.
Host
We talked to Cole who did open Inspect. Yeah. And obviously Real Inspect also is on modal.
Akshat
Yeah. So Ramp Inspect was a great example of a background agent that was really successful because they were able to use some of the primitives like snapshotting and fast scaling to just have something that feels really reactive and works well.
Host
Yeah, that's the new CTO of Ramp right there.
Akshat
Yeah, Rahul,
Host
it was really, really fun. Yeah, I mean, you know, I think all very bullish. Like, you know, one of my reflections was Also, I did not originally, obviously when I met you guys, you weren't that much in a GPU game and now you're all about inference. And one of the points that I hinged on for Jensen's keynote at GTC this year was what we're calling the inference inflection that let's say in AI workloads or machine learning workloads, it used to be like, let's call it 8 to 1, GPU to CPU and now it's more like 1 to 1, which is like an interesting. Because of how much agents basically are blocked or call out to cpu. Heavy stuff. The actual like limiting factor swings back and forth from GPU to CPU a lot more than it used to be. All GPU and then occasional cpu. Yeah, gpu, cpu. And now it's like just constantly and you just have to co locate everything.
Akshat
Yeah, and that's one of the things that actually again we see is something appealing about modal, which is we've built this capacity pool that spans 17 cloud providers. So we're very good at running on various kinds of cloud capacity across the world.
Host
You don't have your own data centers?
Akshat
We don't have our own data centers. We just run across a lot of NEO clouds and metal providers.
Host
Yeah, you're running the math and you're like, what's the cutover point where you're
Akshat
like, yeah, it's a good question. I mean part of it is we see our differentiator in the software layer and being capital light and focusing on the software helps us move really fast. So far it's worked out well because there are so many other people building data centers that we're able to work effectively with them and again focus on what makes us special.
Host
Yeah, 17 gets you into like the local providers sometimes. This was the most interesting one.
Akshat
There are actually a lot more NEO clouds than you expect and they all have various degrees of various levels of reliability. And that's why it's something we've invested a lot of time in is actually building our own reliability layer on top. So if the GPU falls off the bus or something happens, user workloads are not affected. And that actually lets us use a lot more capacity than you as a user would be able to.
Host
It's a useful thing to have because now everyone knows what layer you are and you sort of optimize for being the super cloud of all clouds.
Akshat
Yeah, that's the idea. And so I guess when you mentioned colocation, that's another interesting thing where one thing we've seen is people come to us when they want very specifically located CPUs or GPUs like they want.
Host
Oh, they pin it in EU.
Akshat
Exactly. Or EU.
Host
Data locality thing or performance or what.
Akshat
It's either data locality or latency. They're running sandboxes in modal. They want them to be right next to.
Host
It's easy to. That is important. And all those things you've kind of accidentally. I don't know if it's accident, but you've built the perfect primitive for agents to express themselves. And then it's almost very funny how every extra development just involves more file system, just involves more cpu, just like the things that you already have. I don't know much about if there's any networking usages that are interesting, but you've also done some good work on networking.
Akshat
Yeah, I mean, that's exactly right. We're sort of just taking compute storage and networking and building stuff on that layer for again, the stuff people need. We see a few interesting network things coming up. One is people actually want networked sandboxes
Host
for like a Docker cluster type thing. Sorry, Docker swarm. What is it called?
Akshat
Compose.
Host
Compose type thing.
Akshat
Yeah. So actually if you want Docker compose, our sandbox is now support. Support this thing called sidecars. So a sandbox is actually a pod of containers and you can run multiple containers in the sandbox. Also useful because going back to networking, people want a lot of control over outbound networking from a sandbox. They might want to run a man in the middle proxy for maybe logging stuff for RL or controlling how egress can happen to a domain, injecting credentials. And yeah, so we've kind of had to build a lot of that stuff ourselves. But then also sometimes people actually want sandboxes spanning multiple nodes to talk to each other, which is an emerging thing we're seeing. We have support for that for a different reason. And yeah, we'll see if that becomes safe.
Host
Like just an open socket. This is directly like mtls, we do
Akshat
support that, which is you can expose a tunnel inside a sandbox and then you can either expose the public Internet or it can be. And you can add like a HTTP auth layer above it. But we have this thing called i6pn which we haven't talked about, which is this overlay network using IPv6 addresses. So if modal containers within the same workspace when this enabled can actually address each other using this private IPv6 address and no one else can, this is like sort of private Networking for containers. We actually built it because we needed it as primitive for our distributed training product. So we have this other feature which is you can add a decorator to a function and you get a cluster of GPUs and they have RDMA networking so you can run a distributed training job that's truly serverless. And we did the overlay network for that. But then we've seen that people are using it for other reasons and I'm kind of intrigued to see. Yeah, what would people do with it?
Host
Build primitives and let people figure it out.
Akshat
Right, exactly.
Host
They're like, they read the docs.
Akshat
We think.
Host
Let me use that for something that we never intended to.
Akshat
This is literally not even in our docs page. People somehow found it and they're using it.
Vibu
I mean, the way you portrayed it with like RDMA versus tcp, very well laid out, but just the transfer speed change at scale for rl. Yeah, you have it. You have it built in. I'm sure someone found it, found it to be a lot more efficient before you actually made a thing out of it. Right.
Akshat
Yeah. And not to split hairs, I guess the overlay network actually is the TCP overlay network. The reason we have that is you need that to do the key exchange for RDMA before you set up the RDMA network on top of that. But then people found the TCP part.
Host
Can I tell you, this is like a big aha moment for me because I review 2,200 submissions for the World's Fair and then I got this from John Osterholt, who I don't know if you know John Osterholt.
Akshat
The name sounds similar.
Host
He's a well known professor, published a lot of interesting software design books. And this is the talk he chose to submit. It's on RDMA and inference. And I'm like, you wouldn't think that this guy who is like kind of operating systems guy would care about rdma.
Akshat
I mean, it makes sense to me because. Cloud, right? Yeah. Like the way you move around your KV cache and how efficiently you can do it, how efficiently you move your weights from your training GPUs to your inference GPUs. In RL, there's a lot of degrees of freedom and it is basically a systems problem of moving memory around scheduling.
Host
This shows you how primitive my understanding of networking stuff is. Is this like the domain of wireguard as well?
Akshat
Not quite.
Host
It's adjacent. So explain everything. Sure, sure.
Vibu
How do we move memory around gpu?
Host
Sorry? Yeah, memory. Sorry, I was talking More. And maybe I was talking like five minutes back about the private IPv6 addressing that you've set up. It's basically a VPN.
Akshat
Yeah, it's sort of like a VPN. And yeah, Wireguard is. Yeah, you're right, it is.
Host
Yeah. You already moved on it in topics.
Akshat
A similar. In the same space. Wireguard is encrypted. And this is.
Host
You don't need encrypted.
Akshat
Yes, it's not encrypted. That's the main difference. This is TCP and we have EBPF programs that will reject or allow the TCP connection based on whether you're allowed to do it.
Host
It used to involve a full sidecard, but now you have EBPF in a Linux kernel. Yeah, I don't know if this is a natural. Follow on to the topic of my skepticism on distributed training is that, well, people spend a lot of money on cables to hook up GPUs and even that is not fast enough. And that's the bottleneck. Is your networking fast enough?
Akshat
Yeah. So I guess you're talking about sort of fully distributed training like Dialoco or something, which is like cross. That would be.
Host
Yes, that's the extreme. You're kind of in the middle. And then other people would have like the Mellanox cables up in their actual data center.
Akshat
When you run Multano training on modal, the RDMA, I think Mellanox is, or InfiniBand is like a is. You also use RDMA, but basically it's a way to bypass the TCP networking stack and transfer stuff much faster between one node to the other. And we have I think like 3 terabit per second internal networking, which is the standard that's needed.
Host
Okay, so I misunderstood what part of the stack you were.
Vibu
Tcp.
Host
Okay. Yeah. I mean very impressive work. So effectively you're extending sort of like the modal philosophy to the trading cluster.
Akshat
Yeah, and we're not going for obviously large scale pre training runs. The thing that we've built multi node training for is we see a lot of smaller scale post training. Like people are post training like medium sized quant models so they can get higher quality on inference. This is the perfect fit for something like that.
Host
Yeah, that is my impression of how a lot of these labs explore branches in post training and then eventually merge whatever they find in.
Akshat
Yeah. The other use case we've seen for multi node training is even if you have a big cluster, your researchers are still doing small runs and having elasticity there matters a lot more.
Host
This is Actually the current limiting factor for auto research which is like you basically need to give your model some GPUs.
Akshat
We have a blog post on auto research and modal turns out to be a pretty good substrate for that.
Host
So my impression is auto research means many things. Like the only thing that Andre coins right now. Still science fair, right? Not actually. I don't know how many people are actually doing this.
Akshat
I taught the same thing.
Host
Yeah, you would know.
Akshat
Our internal, both training and inference teams actually use this sort of the general shape of this quite a bit. Like we have this one internal repo called auto inference which is essentially we've automated our own FD efforts using this harness which is the agent will just spin up a sweep of different things. It'll even run like Nvidia inside profiler and it'll like tweak configs and it'll arrive the right thing, it'll change your GPUs. It's from H200 to B200 and it actually works really well.
Host
Nice. By the way, I enjoy that your FDE is so technical that you have to do these things. It's very different from FDE from other people.
Akshat
Yeah. Our FDE team is essentially they're like applied inference researchers or applied training researchers.
Host
Someone told me they have to be able to build, but they also have to be able to sell. Do they have to sell or are they like. They're good. It's just like post sale type of thing.
Akshat
It does. Being able to talk to a customer and engage effectively with them matters a lot.
Host
They all want the same thing, but
Akshat
it's not really a sort of sales thing we pair them with. We have solution targets as well that are more on the pre sales side.
Host
Okay, let's spend a bit more time on auto research. This is a big focus for me for this year. Where does this go? Have people explored enough? There's all these beautiful charts of improve level off a bit and then you find the next thing. Is this basically sort of one abstraction up from normal training. Is that how we think about it or do you think about it differently? Like model level training versus basically AI driven hyperparameter search. Some people call it neural architecture search or whatever. Right?
Vibu
Yeah.
Akshat
I mean so the stuff I've seen people do with it is nowhere on the architecture level. It's pretty much tweaking parameters, but it's basically a hyperparameter sweep that's guided by some sort of model intuition. So it's much more efficient than whatever other sweeper you would have.
Host
Yeah. I mean it's just a question of where you want to spend your computer because you can just throw infinite amounts of money on this and somehow you'll bang on Shakespeare.
Akshat
Infinite Monkey.
Host
Very good for model. And I think it's also very important that agents can spin up. Other agents can spin up their own infrastructure. Very good for you. How good are LLMs at generating modal code? The benefit of existing pre LLMs is that you are in the data.
Akshat
Yeah. They're actually surprisingly good. I think pre cloud 4 they were not and then now they're able to one shot stuff out of the box. We were playing around with releasing a modal bench for the harder things that the LLMs cannot do yet. And maybe what's an example of that? I think the things that sometimes agents struggle with without right. Guidance and a skill is how to use the rest of our observability. Like how to something is failing. How do you look at the logs and then update the right thing? It's sort of reasoning about that. But they're able to. One shot.
Host
Yeah. You can just add a skill to it.
Akshat
Yeah. So we have a modal skill now which is kind of actually why we built this modal bench. It's to find things like that so we can address them in our skill.
Host
Yeah. No, I mean it's good. Are you facing any shortages? We talk a lot about GPU shortages, but also cpu, also memory.
Akshat
We have had a lot of growth which means that there's. We've had to be much better about proactive capacity planning.
Host
Yeah. Which by the way, it's like an MBA's dream job. It's just planning this stuff. I think last time you and I talked about something about this.
Akshat
Yeah. I mean we have a really competent team of people. The role is called compute strategy. So yeah. If anyone listening here wants to work.
Host
Compute strategy.
Akshat
Yeah.
Host
I mean the normies call it FP&A or something.
Akshat
Well, it's not FP&A. There's a lot of interesting financial questions of what is the blend between one year and three reservations. How do we forecast our own capacity? How do we basically. Especially since our capacity is very fungible across different GPU types and different regions, you basically have to model a lot of it and you also have to have an opinion on how the supply chain is going to evolve and. And then you have to take bets
Host
based on that tokenomics. This is probably not a real point, but I was trying to think about what other industries I always try to think about. We cannot be first to These kinds of problems and what other industries have had this. And I was like airlines with fuel and they have to hedge their fuel. And I think for a long time Southwest because they made a hero fuel bet they were super low cost because compared to everyone else.
Akshat
Yeah, I hadn't thought about that.
Vibu
We're at a fun time too, you know.
Akshat
Yeah. A lot of the compute business in general for us is also about being very good about capacity management. That is how you have great unit economics but also over time it's how you can unlock more value for customers. Like one of the things we're building now is a way for customers to get, if they don't care about latency, like get much cheaper pricing and they'll get results back in next 24 hours. Something like a batch tier essentially.
Host
Yeah.
Akshat
And those are levers we have because we control the whole stack and scheduling and whatnot to give people a sufficient.
Host
Yeah. I feel like they're not as popular. The Frontier Labs have all those APIs. They're not as popular as they should be.
Akshat
The demand that we see for something like that is actually not for LLMs. Although sometimes people want to run evals and do synthetic data prep and there it makes sense. But it's from a lot of non LLM companies. Like people who are doing computational bio. They haven't run really big batch odds and they don't care about when they get it back.
Host
Yeah. And they have a reasonable. It's also like a cousin to the stopping problem of like will this finish in time?
Vibu
Yeah.
Akshat
You can bound it, you can give people SLAs on it. Yeah.
Host
I think what's interesting is the next phase of modal. What do people expect from you now that you're sort of established and you're a well known computer player among all these leading companies. You had an inference launch week and we talked a little bit about the launches. What else? What else should people know?
Akshat
We are building primitives that make our users lives much easier. So I think, for example LM inference. Thousands more companies are going to post train their own models and deploy open source models for inference. So we're thinking a lot about what is the best product shape for that. And that involves everything from our training gym to then endpoints that get frontier level performance again without having to talk to anyone. It looks somewhat different on other verticals. We're also seeing a lot of real time audio, video stuff in there which is why we're working on things like regional routing with fallbacks so you can get sort of GPUs that are as close to users as possible, so you get like low latency for video streaming and whatnot. And then on the agent side, we're still working very closely with our customers because stuff is changing so fast in terms of what they need. And I think beyond sandboxes and persistent file systems, there's a lot of other things people need from this agent stack as they build production agents. So yeah, we're thinking about those other things that fit in there.
Host
I want to ask what the other things are.
Akshat
Yeah, I probably should share right now, I think.
Host
Okay. So I do think a lot about the principal components of cloud. And you do talk about compute, storage, networking. Because so far for me it's fine. So far for, I mean the first couple generations of cloud, it's fine. What's different, qualitatively different about agents that you need some new permission level? A lot of people obviously.
Akshat
Okay.
Host
And I'll just kind of spew tokens at you until it hopefully sparks something.
Akshat
Yeah.
Host
The new level now is whatever cloud code does, which is dangerously secure permissions or allow this by command or whatever. Right. And sometimes they're like, okay, well we have this adaptive thinking mode where just trust me, bro, I will make the calls for you. Is that it? You know, like basically like sort of LL mediated permissions.
Vibu
Now you're looping it with a goal and letting it roll.
Akshat
Yeah, I mean I'm, I'm skeptical of LM media permission for stuff that is at the sandbox level because you do want hard boundaries. Yeah. Otherwise obviously someone can exfiltrate stuff.
Host
But like, like maybe, maybe that's old school thinking. Maybe we're the dinosaurs. Maybe the AIOS or the LLM OS is really. The kernel is a goddamn LLM. It makes you feel uncomfortable.
Akshat
Yeah.
Host
But that's what trusting the LLM is. Imagine a spherical cow. Perfect LLM.
Akshat
Right.
Host
Let it. Maybe I want to test the boundaries. Right. Obviously I don't believe that, but I want to see where I'm wrong because that's the non consensus.
Vibu
Yeah.
Akshat
I mean I think you always need hard guardrails when you want and you can pair those with software guardrails. Right. And that's gonna be a lot
Host
end with a couple of your commentary on like the ecosystem outside of modal manage agents, everyone has 1. Gemini OpenAI Claude. Very useful for you. But also like it is their way of starting to edge into your space. Yeah. What's going on?
Akshat
Yeah, I mean we're very excited to partner with anthropic and some of the other foundation labs will not name who we're also working with. The way we see it is the managed agent thing is a great place to start if you're starting out building an agent. But then when you get to building something more production grade, like you're a company that's like ramp, that's building their own. Ramp also runs their accounting agent on us, so their external facing agent. You need a lot more control over your compute primitive on things like what sort of how do you process different files that the agent has access to and how do you snapshot and restore. How do you control the networking? Maybe you want GPUs. When you get to that point, you kind of want a specialized sandbox provider that gives you those things. And that's the role that we are trying to play. We don't really have an opinion on the harness, whether it runs in cloud managed agent and you hook it up to modal sandbox or you run the harness in modal sandbox. We'll see where people converge with that.
Host
Yeah. Do you have any opinions on the meta harnesses? It's just another layer on top of these things.
Akshat
You mean like the OpenPi?
Host
OpenPi is one. I think Vercel had one, which I can't remember the name of right now. Fred Schott had one and then to me most recently was data bricks that had omnigen. All these are sort of meta hardest like kind of pseudo agent cloud type things.
Akshat
I personally have not played around with them.
Host
I mean everything's bullish modal as long as it consumes more infra.
Akshat
That's why we're focusing on the infra layer. It's somewhere where our relative competences and also it's a hard problem to solve.
Host
Yeah, I mean I will say just generally reflecting on. I don't know if there's other topics on model, but just generally reflecting as an infra person, not as intense as you, but in that field this has been the most exciting time in infra. It was boring actually for a while and you couldn't really get people excited about data infrastructure. Eric would get on data console. Everyone just watched the video and look at how many sandboxes I can spin up and no one gave a crap. Yeah, and now everyone gives a crap.
Akshat
That's true. It is a very exciting time and I think a lot of that's driven by just the amount of scale all of this stuff needs.
Host
I think a lot of your initiatives, a lot of your product directions make sense in retrospect. Which is the best kind of. But I wouldn't necessarily have thought about it myself.
Vibu
We need the predictions. I mean, I think there's a lot that you just don't even see. Right. Like you have the batch, you have the voice, you have the multimodal. But what else, you know, what else
Akshat
is coming up for us?
Vibu
Where do you see things going?
Akshat
Yeah, I mean in general it's clear that there's obviously there's a huge shift happening. I think one thing that's not as obvious to people because LLM inference gets talked about so much is also we work a lot of companies that are doing things like drug discovery and computational bio, like the China discoveries world. Big things are probably going to happen there. We work with a lot of robotics companies that are actually putting robots in active deployments and getting good results out of them.
Host
Is there air gap model? Is there a version that is like on prem. Air gapped, whatever?
Akshat
No, we usually cloud only.
Host
Yeah. Okay. But yeah, I mean, so what you're saying is because you're focus on primitives and they're good primitives, you find use cases and all these kinds of things actually probably diversifies you a little bit away from LLMs all the time.
Akshat
Yeah, absolutely. Our goal isn't to only serve the LM and FIST market.
Vibu
Just on the website, the audio, the push up. We've had bolts on the Bino images. Yeah. I mean there's a lot here. There's QTI tts customizing. Oh, Chatterbox. There was a customizing whisper.
Host
This screen reminds me of a fallen competitor which replicate. What's your post mortem on what happened?
Akshat
This is one thing we've kind of stayed away from is providing an API for models. Because I think providing model APIs is some of it ends up serving like a really hobbyist market which is much less sticky. And we've always wanted to build for companies that are building sort of products and need sort of more flexibility. That's not just an API which you
Host
can build an API for a model and this is clearly what it is. But you're saying you can wrap it into a more fully functioning backend that you run.
Akshat
Yeah. So actually all of our examples, it's not that spin up this model. Here's an API token, use it. They're actually all code.
Host
Okay.
Akshat
And so the point is that this is an example starter code. Yeah. But you can tweak it however you want. And if you're like a company building a product like computational bio whatnot.
Host
Yeah, I guess I'm Trying to tease out for listeners when does it stop becoming oh, you're just an API call and you're just a wrapper on an API to becoming what you call a product. Right. Like what is that layer? Obviously more lines of code but beyond that what is the substance that people add that qualifies it to be something more?
Akshat
I think there's a little bit of a selection effect of a lot of companies who do want to get deeper into that level are probably building something that's more differentiated. I think an example is with LLM inference. Originally we worked with companies that were building their own post training frameworks or they were ramp actually early in the day was training their own tokenizer and swapping out the tokenizer in llama and whatnot. I'm not saying that successful in that case. A better example is let's say suno because SUNO does not use modal for
Host
training Mikey on the pod.
Akshat
Yeah but they use modal for all their inference and that's because they have a custom, they have completely custom model architecture and that means that they have to be at the code level and tweak things that are not. Yeah, it's an API.
Host
It's interesting as well. We had Ethan most recently on the XAI Grok team make a prediction that actually the next tier in videogen is not a better video model. It's a better model or agent that orchestrates video models, language model backboard that
Vibu
can use tools and write code like
Host
yes, I can make my 6 second video or my 10 second video from Grok but actually I want my 6 minute video and I'm not going there through normal video. Jen.
Akshat
Yeah, that's interesting. We have GPU sandboxes and recently have seen a few companies doing sort of agents that do video manipulation or yeah,
Host
give it FFMPEG and just run ffmpeg. It's not enough. You need to give it Adobe.
Akshat
Yeah, I hadn't put it together with like it would actually be a video production thing in my mind these things were going more towards editing and yeah,
Host
I think about this a lot. Obviously.
Vibu
Sorry Luma. Luma Agent is a version of this for video production but you know it's a one off.
Host
I was going to get your quick takes on some other stuff that happens in recent news and see if you have anything interesting. Gitpod very somewhat different market. They're in the CI CD market but actually technically very impressive. I don't know if you've taken a real look at them.
Akshat
Yeah, people on our team have talked to the gitpod team and they're technically very strong. We're very bullish and modal on the CI market as well because there's more agents, coding agents, they're going to run a lot more CI and the primitives there can be much better.
Host
I think there's a lot of wasted CI.
Akshat
Yeah.
Host
So is it just like let's filter what is the highest order bit here in improving CI for agents?
Akshat
Well, there's a lot of wasted time in CI on preparing your artifacts and getting you to the basically preparing your dependencies and whatnot. And obviously build systems help with that. But if you have primitives that are memory, snapshot and restore, can you just run CI more efficiently?
Host
Oh, okay. Interesting. Yeah. I mean another form of on demand compute.
Akshat
Yeah, exactly. It needs the same again platform
Host
for those who don't know. Gitpod rebranded to ona. It was like there was this whole thing actually semi sounded the alarm recognition. I was like, you should take these guys seriously because they're infrared. Very good. But then they joined OpenAI and presumably we'll see codecs cloud from the owner team, which I think would be very, very strong to me, teams like that that can set up the networking and the secure boundaries for your agents to have their own cloud. Each effectively is what you're doing kind of. And I'm just trying to draw the analogy or the differences. If you have studied them. What is the philosophical difference?
Akshat
My sense is maybe they didn't go off to the right market at the right time because I got lucky with like agentic use cases really taking off and dealing more of a sandbox shape thing than mindset.
Host
Sandboxes work like CRCD is sandboxes. It's just like build time sandboxes versus runtime sandboxes. And actually it turned out runtime was better.
Akshat
Right. The difference there is runtime sandboxes have a different configuration surface of the how you configure images, how you attach storage.
Host
Yeah, it's fascinating. Other people, Astro also OpenAI also like Python tooling ecosystem people are you still sort of bullish building on top of Python? Also recently Modular also got bought by Qualcomm. Just any of your takes there?
Akshat
Yeah, I mean we had Python as our first SDK language because that was the language that people did data and ML in. I actually now have Go and TypeScript SDKs as well. And our runtime is completely language. It isn't in rust, but it's not tied to Python by any means. We haven't seen, I think with like inference and training stuff, people are still very Python. And the interesting thing with like the agent stuff is people use our TypeScript SDK a lot more because they're not actually doing anything that needs ML. I don't think we'll have to go beyond that super soon because Python and TypeScript is still dominant.
Host
The last two languages in the world. Yeah, that's it.
Akshat
Well, English and prompting is the English and prompting.
Host
I occasionally talk to people who try to build new languages. They're like, even what's his name? Brett Taylor, who's chairman of OpenAI, was like, we need a new language for LLM. So no one has come across one. And I keep looking, you know, Python and Typescript, you have a lot of data plus. But then also they are very imperfect as just as languages themselves. Then my close is, I think modal used to be a big bet on developer experience and you've pivoted the team to agent experience. Is it like the way now, like, can entire companies and unicorns, multi unicorns, be built on just having better agent experience? Da. Something else.
Akshat
It's a big part of our identity. It's not just the very tactical how does an agent use the cli, but it's also how easy is it to spin something up? What is your iteration time when you want to spin up a new service and you want to get something going in prod, in practice, that matters a lot to people and I think it will continue to matter. People are building stuff even faster and if you give them ways to do it quickly and not have overhead, then
Host
I think the debate for me has been do you do anything differently that is very fundamentally different for developer experience versus agent experience. You seem to be on the side of they're like this.
Akshat
I actually also have a blog post on that.
Host
Cosine similarity on 0.9 or whatever.
Akshat
Yeah, I mean, pretty much the main shift for us has been, as I said, we built this, this benchmark modal bench to see where agents are lacking and actually literally add surface areas to a product. If they're reaching for something, maybe this should just be a cli.
Host
Yeah, they hallucinate their own features.
Akshat
Yeah. And sometimes it makes sense. If they're reaching for this thing, it's product feedback, give it to them. And then. Yeah, actually moving. We used to only have logs and metrics in our ui. Just moving all those things to CLI as well, so they're accessible in that form.
Host
Simple as that. Cool. Thank you so much. Yeah, great update. And I can see why you guys have succeeded so much. It is really focused, but also really good execution.
Akshat
Thanks. I mean, we have a long way to go.
Vibu
All right.
Host
Thank you.
Akshat
Co.
Episode Title: Why AI Infrastructure Must Evolve for Agent Experience
Podcast: Latent Space: The AI Engineer Podcast
Guest: Akshat Bubna, CTO of Modal
Date: July 8, 2026
This episode dives deep into the evolving needs of AI infrastructure, focusing on how agent-driven workflows demand new primitives and experiences. Akshat Bubna traces Modal’s journey from a developer-focused serverless platform to the backbone of production AI applications, unpacks technical innovations in bursty, elastic compute, and discusses how Modal’s approach to observability, auto-scaling, and “agent experience” positions it at the forefront of the AI infra landscape. The conversation covers real-world use cases, technical details, and big-picture trends shaping the next generation of cloud and agent-native platforms.
On Collocating Infra With Code:
“You could really condense the surface area of what you're doing, put in code so you can actually operate on it just like you can operate on other code and build stuff that's more expressive and dynamic.”
— Akshat, [03:40]
On the Rise of Agent Experience:
“We've actually changed our SDK team to think about agent experience...the same benefits that apply for DX also actually apply for AX.”
— Akshat, [04:54]
On Auto-scaling and Elasticity:
“We've incorporated GPU snapshotting into the product so we can actually take the GPU state...and next cold starts are way faster.”
— Akshat, [13:36]
On Open Source Innovation:
“By using open source DFlash you can get the same performance as with proprietary providers. The next thing we're thinking about is helping people train speculators and custom models...to make frontier-level performance available to everyone.”
— Akshat, [16:08/18:08]
On Capacity & Infrastructure Management:
“Being capital light and focusing on the software helps us move really fast. So far it's worked out well...there are so many people building data centers that we're able to work effectively with them and focus on what makes us special.”
— Akshat, [24:35]
On Target Audience:
“We've always wanted to build for companies that are building products and need sort of more flexibility... It's not just an API.”
— Akshat, [48:58]
On DX vs AX:
“The main shift for us has been...we built this benchmark modal bench to see where agents are lacking and actually literally add surface areas to a product. If they're reaching for something, maybe this should just be a CLI.”
— Akshat, [56:53]
The episode expertly charts Modal’s journey and strategic bets as underlying AI infrastructure moves from developer-centric, human-configured “cloud 2.0” to truly agent-expressive, burstable, real-time, region-aware “cloud 3.0.” Akshat’s thesis is that both hardware and software must adapt in lockstep—with agents, not humans, now driving the majority of scale and shape of workloads. Modal aims to serve as the "super cloud" of agent infrastructure, with open-source innovation, rapid elasticity, and minimal friction for both startups and advanced research use cases.
If you want a technical but pragmatic lens on where the infra layer is going for AI—and how that impacts both agent developers and production enterprises—this episode delivers both the why and the how, with hands-on examples and honest reflections.