![Etched - Building AI Hardware to Make Inference Faster and Cheaper - [Invest Like the Best, EP.480] — Invest Like the Best with Patrick O'Shaughnessy cover](https://megaphone.imgix.net/podcasts/532910c4-743d-11f1-b7dc-b31915c88b79/image/810c9e3a23b5c4ea0a56e5e4cc8d7e47.jpg?ixlib=rails-4.3.1&max-w=3000&max-h=3000&fit=crop&auto=format,compress)
Loading summary
Patrick O'Shaughnessy
I know firsthand how complex the tech stack is for asset managers, and seemingly every new tool and data source makes the problem even worse, adding more complexity, more headcount and more risk. Ridgeline offers a better way forward, one unified platform that automates away all that complexity across portfolio accounting, reconciliation, reporting, trading, compliance, and more. All at scale Ridgeline is revolutionizing investment management, helping ambitious firms scale faster, operate smarter and stay ahead of the curve. See what Ridgeline can unlock for your firm. Schedule a demo@ridgelineapps.com OpenAI cursor, anthropic, perplexity and Vercel all have something in common. They all use work os and here's why. To achieve enterprise adoption at scale, you have to deliver on core capabilities like sso, scim, RBAC and audit logs. That's where work OS comes in. Instead of spending months building these mission critical capabilities yourself, you can just use WorkOS APIs to gain all of them on day zero. That's why so many of the top AI teams you hear about already run on WorkOS. WorkOS is the fastest way to become enterprise ready and stay focused on what matters most, your product. Visit workos.com to get started. Felix Byrogo is a personal finance agent that turns a single prompt into finished client ready work using your firm's own templates, context and standards. Send Felix an email like Take these comments and turn them for me or update my tracker with the context of these emails. Or run the ability to pay math on this buyer and Felix sends back finished PowerPoint decks, Excel models and sourced research. Felix works the way your team already does, delivering work quickly and accurately around the clock. Learn more at rogo AI/felix hello and welcome everyone. I'm Patrick O' Shaughnessy and this is Invest like the Best. This show is an open ended exploration of markets, ideas, stories and strategies that will help you better invest both your time and your money. If you enjoy these conversations and want to go deeper, check out Colossus, our quarterly publication with in depth profiles of the people shaping business and investing. You can find Colossus along with all of our podcasts@colossus.com Patrick O' Shaughnessy is the CEO of Positive Sum. All opinions expressed by Patrick and podcast
Gavin Uberti
guests are solely their own opinions and do not reflect the opinion of Positive Sum. This podcast is for informational purposes only
Patrick O'Shaughnessy
and should not be relied upon as
Gavin Uberti
a basis for investment decisions. Clients of Positive Sum may maintain positions in the securities discussed in this podcast. To learn more, visit PSum VC.
Patrick O'Shaughnessy
My guests today are Gavin Uberti and Rob Locken, the founders of Etched. A few years ago when they set out to build a better AI chip than the largest companies in the world, almost everyone I called told me it could not be done. They've since done it, taping out a working chip on their first attempt and becoming the first hardware company founded after ChatGPT to do so. They already have more than a billion dollars of customer demand for their first product and have raised $800 million to build it. Etched builds chips and systems designed to run AI models faster and at lower cost. They started the company in 2023 and the product is a complete rack for Inference. The chip along with the boards, the power delivery, the interconnects and all the manufacturing to produce it all. We talk about the technical bets behind their architecture, how they hired industry legends and paired them with elite 22 year olds, and why they believe Inference will become one of the largest markets in the world. I think you will find the story of what they have built hard to forget. Please enjoy my conversation with Gavin and Rob. All right, gentlemen. It's been three years or so, Gavin, since you and I last did this, which is nuts. And at the time I was just wildly intrigued by your story and what you were going to build. I didn't know a lot about chips at the time. I was considering investing in the company and so I was calling everyone I couldn't conceive of that could give me an opinion or something. And at the time, basically the consensus was these kinds of companies are not built by young people, that the semis world, the best companies are founded by 40, 50 year old people that have had a whole career's worth of experience, have learned all the problems, have shipped multiple chips. 221 year olds are like not going to do this. It's just not going to work. It was indicative of a theme which was nobody believes in us. That's obviously changed a lot now. You know, just walk the halls and talk to the people that have chosen to come work here. But in the early days it felt like this was something that you had to face down. What was that like facing that down where a set of incumbents and an industry worth of people and investors and everyone else sort of didn't believe in you. And what did that anneal in you to build the company the way that you are? Like, what was the impact of that?
Rob Locken
I think there's a certain level of naivety required to think that you could build a chip better than every other AI chip ever built and build A company to do it way faster than ever has been done. And we have the naivety, but there's many times where we would say like, why isn't this possible? And really push on it. And it turns out that everybody's answers are extremely siloed to a set of constraints that aren't true anymore. And the reality is the entire semiconductors and data center industry is built on buffer. And what I mean by that is every part of the stack, from the EDA tools to the power modules, to the circuit boards, to the chip design and standard cells, everything is built to be general purpose for everything. Not just in the data center, but Iot on the edge and so forth. And when you have a specific use case you're really trying to design for, you can change the constraints a lot. And I'll give you just a very simple example that we're not the only one who does, which is one of the things you care a lot about is the clock speed of your chip. It's proportional to the throughput of your system. When you are doing sign off for different timing, what clock speed you're actually going to be able to run on. When you tape out your chip, there's this concept called corners, which is, you know, what temperatures are you going to be able to run at this clock speed. The default configurations for a lot of these EDA tools assume that you're going to be running your chips in freezing temperatures. Now, I don't know about you, but I've never seen an AI data center with ice in it. So, you know, we can feel pretty confident that our chips don't need to run at full speed at 0 degrees Celsius. In fact, like they're never really going to be running below 80 degrees Celsius anyway. And just by knowing that that's a constraint that doesn't matter, we can make a ton of changes throughout the entire system. That's like a very simple one, but there's many more that you get 20% here, 50% there, 2x here. And these compound to a system that can be radically better for inference.
Gavin Uberti
I think you found two kinds of people. There were some folks who went purely on heuristics of hey, young founders. They claim they can go beat the biggest company in the world on performance. It cannot happen. And there is no thing you could go say to me that would make me change my mind. But there's also people out there who are of course skeptical but are willing to go ahead and say, I'll spend the time, I'll do the work. And is it actually possible? Like for example, one of our earliest, earliest supporters was Mark Ross. And Mark was a very prestigious semiconductor expert, used to be CTO at Cypress Semi that sold for $9 billion. And when we met him, we were just a couple of guys in a dorm room. And we came to him and said, hey, you want to go build hardware for inference, meaning we can be much faster than Nvidia. And Mark's like, no, you can't, it will not work. But if you want to go convince me, you should write a white paper. You should go ahead and build a functional simulation and show me. And so after a lot of very long nights, went back to Mark and said, hey, here's a simulation. What do you think? And he was like, huh, this works. But to go do a company like this, you'll need a large amount of capital and at least $3 million even to get up, get started, end up, went ahead and raised five, raised a lot more after that. And then he was again surprised, but got more involved. And then he became an advisor and a halftime advisor and eventually a full time CTO as he saw more and more of the development progress. And I think in general this is filtered really heavily. Folks who want to go ahead and be right regardless, folks who want to be very truth seeking and say, sure, I'm skeptical, but I will go ahead and work through the numbers myself. And if I can go figure out why this is possible, well, let's go build it.
Patrick O'Shaughnessy
The specifics that you've made bets on the way that you built this system are immensely interesting to me. And because so many people are trying to do this now, build new chips that will do a better job of serving inference at massive scale. The world is interested in the research approaches, the different architecture approaches that people are taking to building a new AI chip. And I'd love you to just start by describing what this thing is, what it does, but maybe more interestingly and more importantly, the process that you went through to decide what bets to take, what technologies to invent and compare and contrast those with what you've seen the rest of the marketplace try to do.
Gavin Uberti
Yeah, I think that you can start with the product. We're not just building a chip, we're building a full inference solution. And that means a rack. That means the chip, that means the power delivery into the chip, that means the board on which it sits, that means the interconnect by which the chips talk to each other. That means the production for this mass volume of racks, really, the production Is the product. We think about how we get our advantage. There are two key parts of running inference. There's pre fill and there's decode. Now we have two key type matching both of these things. Prefill is reading in a huge volume of text and decode is then using that data to generate output tokens. When you go out and run prefill, your key job is not to go predict tokens. You already know the text. Your job is to go and get the model's memory, what we call this KV cache, into the write state. Then you can go ahead and run decode with that same KV cache. So we will often go do is we call PD disaggregation prefilled decodedisag. You will have one cluster of servers running these pre fills. You'll then transfer those model memories, those KV caches over to the decode cluster and then go ahead and use that cluster to go generate the next tokens.
Patrick O'Shaughnessy
So sort of like loading the gun and then firing it. Like if I think about it in super simple terms.
Gavin Uberti
Yeah, you got it. It's getting the model to remember the right things and then using those things to go do tasks.
Rob Locken
Candidly, people think about this market a bit lazily. They say, are you prefill chip? Are you a decode chip? If you're a decode chip, are you an HPM chip? Are you an SRAM chip? Are you 3D DRAM chip, are you using optics, using copper? When we started this, we just wanted to understand why extremely smart people were working on these different directions. We seriously looked at architectures like having a bunch of DDR memory and like a shared memory pool and looking at advanced packaging to basically break out of the shoreline. We looked at things like the is there ways to put memory dies on top of compute dies? In doing so, we realized that there's no free lunch. Everything has a trade off, right? 3D dram, you have a thermal issue, you have a supply chain issue, you have to figure out hybrid bonding, you have to figure out the flops. So now you're a decode chip. So we went through everything both on the prefold and the decode side. In doing so, we realized there's a few design spaces that nobody had seriously tried to explore because they were never done in AI chips. And we asked ourselves what are the actual metrics that are going to matter the most on the pre fill side? The thing that matters is flops and flops density. And people talk about flops often as a headline number, but in reality you should care about the flops you're getting. When you're running real workloads. There's this concept called MFU or model flops utilization, which is, you know, for every peak flop advertised, how many cents on the dollar are you actually getting? And on GPUs you often get somewhere between 20 and 50% depending on the workload. And actually you can provably not run at 100% because you have a thermal issue where as you increase the flop utilization, you have more transistors going on and off, you draw more power, and the chip will self regulate and actually lower its clock speed to make sure it doesn't overheat. So as we looked at inference, we said if we want way more flops, because we want to run at way higher throughputs, we fundamentally need to solve the thermal problem before we even think about adding flops to the chip. If I just add more flops to a GPU today or another AI chip, I'm not actually going to get more performance because it's just going to thermal throttle. So fundamentally the essence of that is this concept of Dennard scaling, which is voltage is quadratically proportional to power. So if I 2x my voltage, my power goes up by 4x. If I cut my voltage in half, my power goes down by quarter. So we asked ourselves, how could we run voltages lower than GPUs? And we talked to a lot of people about this. We flew out to Silicon Valley after dropping out and basically asked, you know, dozens of people in semiconductors at all these different chip companies how they did it. And the answer we got was like, you can't, you can't run at voltages lower than GPUs. And this was very dissatisfying because there was many different industries of chips that run at voltages lower than GPUs. Bitcoin miners run at under a quarter of the voltage of GPUs. So this is obviously physically possible. The question is, are there issues with GPU architectures that make it unable to run at these voltages? And when we looked at the problem for a long time, we were able to create a new mechanism of running at much lower voltages, a new type of power delivery that we call low voltage inference. And we think all AI chips in the future are going to be low voltage chips. They're going to have to cram way more flops in the same silicon area and without thermal throttling run at way lower voltages.
Gavin Uberti
So that's pretty simple for decode, it is All a memory gain, more memory bandwidth, you can load the model faster, load the KB cache faster, and serve more tokens per second per user. We think people ask the wrong question here. People often ask how much memory bandwidth is on your chip. It should be asking how much memory bandwidth is on your full scale up cluster. What we were able to do is add way, way more bandwidth and a much lower latency from chip to chip to our interconnects. It allows us to be able to go serve models at this much higher speed because you can go use the SRAM and the HBM from the full scale up cluster as a single pool. And that's our second key technical bet, what we call cluster scale memory. And on GPUs today, the cluster memory balance is often very badly utilized because the time to go hop from one GPU to another is extremely long. For example, on Blackwell chips, it can be about 4,000 nanoseconds to go point to point. And that means that if you go ahead and go to an 8x TP setup, you will get way, way less than an 8x improvement in your tokens per second per user. And what we did is built our own totally custom interconnect stack. We took everything above the second layer, ethernet, built it full custom, and we can go out and do far, far better latencies and bandwidths this way too. We can go ahead and cut this by more than a factor of 5x. And that allows us to then use the memory of other chips much more effectively. As you scale the world size, your time per token goes down proportionately.
Rob Locken
Yeah, and it's not that surprising given all these architectures were built before ChatGPT. So if we're trying to build a chip for the modern workloads, it's going to look very different. The way we, you know, organize our flops, the way we do our voltage domains, the way we do our power planes are going to look super different. The way we do the packaging is going to look different. The way we do the board design is going to look different, and then the decode side, the way we connect, everything is going to look very different. So we're now bringing forward our first generation of this low voltage inference technology, which is running at under half the voltage of any other AI chip.
Patrick O'Shaughnessy
If you zoom all the way out, why is this so important? Like, why is the delivery of much higher throughput, much lower cost per token, better tokens per watt, like all of these metrics that the universe is going to start talking about more and more and everyone knows the supply side of the equation is a big problem right now. Why is this in a bigger picture, looking at a decade, the bottleneck in the technology world?
Rob Locken
Well, I think it comes down to productivity. Where we are at this extremely interesting moment in the history of civilization where there's real artificial intelligence, not like sci fi stuff, but these models can solve problems that most humans can't. And it's going to create new scientific discoveries, it's going to create instant access to medical care, instant access to education. And now it's just about how many people can use this at the same time, how many products can serve this at the same time, and also the speed of doing different tasks. So when you think about wall clock time, if we can take an agent that can run at a certain model quality and could take a year to solve a certain task using inference time compute, if you have way faster decode speed, you can compress that into a month. So the amount of scientific innovation and the amount of actual proliferation of technology will happen much faster. And then the second part is concurrency, where today it's just not possible for a billion people to use these models concurrently. Ultimately some people are going to get downgraded, some people's models are going to be slower, some people just won't be able to access the hardware. A few years from now there's going to be giant models serving billions of users. We're very much in the early innings of AI today where the paid plans. There's only a few million users in the world using paid plans of AI models. So we're at 1, 1000th of the global population actually using this stuff. So if you want to serve at giant scale, a lot of things change. And one of them is the number of chips that communicate together. Where people usually think about this in the context of training. You have these giant training clusters, you have colossus with over 100,000 GPUs that are all networked together. And the inference side today people usually think about it as an 8 chip cluster or maybe just NVL 72, the scale up domain. But very quickly this is going to become thousands of chips and tens of thousands of chips. And the way to get the most performance there, the time between sending data from one chip to another. That primitive matters way more than is getting credit right now. So when we think about optimizing memory bandwidth for the system, you have to think about how fast these chips can communicate together. Because if they can only communicate really quickly with themselves and very slowly with other chips, you're not going to actually be able to serve giant models at 10,000, 20,000 tokens per second. So we need multiple orders of magnitude of infrastructure built out throughout the entire stack from the wafer to the watt transistor to the token to actually bring this stuff to the world.
Gavin Uberti
And I think that you look at most other goods like the iPhone for example, they've gotten this economies of scale where as a result more money does not really buy a better iPhone that if you're a billionaire or if you're just the average American, you buy the same phone. And tokens aren't like that yet. We're still in the very early days where a general purpose system, relatively small one, is kind of handcrafting each tokens like they made screws back in the late renaissance. And I wanted to live in a world where you have the same economies of scale for token making that you do for making say iPhones or cars or anything else. I think that is one of the huge unlocks that allows a huge group of people to go use the best quality models. Economies of scale have all made capitalism very, I don't know, fair. I think that allows you to go ahead and have the same product in many, many different hands and you're able to go then serve way more users on a single scale up cluster allows you to get closer to that point for token serving too.
Rob Locken
Yeah, and also just certain products aren't usable if they're slow. So if you want to serve coding models and you want people to actually use them, there's a certain number of tokens per second you need to hit. So the question is, while maintaining that per token speed, how many users can I serve at the same time? And you can basically decide I'm going to shut off a bunch of the world from using this stuff or everyone's going to get a worse experience. So fundamentally you need to find ways to push out the curve and that's why there's such a pressure for new hardware.
Patrick O'Shaughnessy
Vanta automates security and compliance for over 16,000 fast moving companies like Ramp Cursor and Harvey, keeping them audit ready around the clock. It's the number one agentic trust platform and it now helps companies like yours watch for the risks that show up between audits across your vendors, your AI tools and your whole environment. Every new tool your team signs up for, every vendor that turns on AI features is an opportunity for something to go wrong. And most security programs weren't built for AI's pace of growth. The Vanta agent works like a 247 GRC engineer in the background, finding issues, drafting fixes for you, and cutting vendor assessment time by up to 50%. Whether you're a fast growing startup or a global enterprise, Vanta helps you earn and prove trust. Invest like the best listeners. Get a Special offer for $1,000 off@vanta.com invest Ridgeline is the first end to end system of record with embedded AI for investment management firms running portfolio accounting, reconciliation, reporting, trading and compliance on one unified platform. Firms are moving off legacy technology and onto Ridgeline because of how far ahead Ridgeline's AI features are compared to anything else in investment management software. I've been hearing from a lot of investment managers about AI and they fall roughly into two camps, with some unsure of where to even start and others convinced they can build their own order management system over a weekend. The reality is that running an investment firm will always require governance controls and a single source of truth for your data, and no amount of AI enthusiasm changes that requirement. Ridgeline is built on exactly that foundation, which is why I believe that firms that come out ahead in the AI era will be the ones running on Ridgeline's unified platform. If you're serious about your firm's AI strategy, Ridgeline should be part of that conversation. You can request a demo at Ridgeline AI. I'd like to take some time to step back and hear both of your stories for how you came to this idea in this company and then kind of walk through what it's been like to build it. Because I think in so doing we'll understand the system that you've built for the company itself that will then be able to power subsequent generations of products like this one. For this crazy inference feature that we're staring down. Rob, maybe starting with you, just take it however far back you want. But what I'm curious about and your personal story was actually the very first thing I ever heard from either one of you was your personal story many years ago now, which really kind of blew me away. I'm most interested in your motivation ultimately for being here doing this thing.
Rob Locken
It starts back actually in high school for me. I've been very unlucky and lucky at different points in life. This was one of the tougher times. At the end of my sophomore year of high school, I got injured at a martial arts tournament the next day. Couldn't walk for some reason. I thought it was something wrong with my SI joint or something. I went through physical therapy. I did different types of scans. I couldn't figure it out and eventually they Found this big bump on my back and an MRI and told me it was a tumor. Stage four bone cancer. Was told I had under 30% chance of survival. It was a two year crazy chemotherapy surgery, learning to walk again, Experience. And when you go through something like that, it changes the overton window of human experience and makes you appreciate what actually matters. And you also ask yourself, what are you going to do if you have the chance to live? If you actually want to get through something like that, you need to be hoping for something. And I always knew I wanted to do something very impactful if I had the chance to get through it. And it took me a couple years to figure out what that was going to be. And at the same time as I got to college and met a bunch of other people building cool tech, I got extremely excited by AI models, especially once GPT3 came out. And I was like, wow, this is the first model that can kind of speak English and these things are going to get really smart. What happened was when GPT4 came out, there was GPT4V, which was the first model with image uploading. So I went through my camera roll and I found a picture of my back with this bump on it before I was diagnosed. And I said, hey chatgpt, pretend you're an expert doctor. A patient comes in and it says they have this bump on their back. What could it be? And it immediately says, this could be a tumor. You should get an MRI immediately go to the doctor. And I just kind of sat there still. It was like, that took me six months and yesterday this feature wasn't there. Today it's here I go to show my parents. And I got this notification being like, you're all out of image credits today, you need to get a pro plan. And I was like, holy crap, this is going to change everything. And we clearly don't have the infrastructure to serve it. There's very few things you can work on that can actually bring this technology and at scale to the world faster. I mean, there's like plenty of people that are super smart working on models. The fabs seem maybe unreachable to work on, but it seemed like the hardware was all designed before ChatGPT. Every GPU, every TPU, every AI chip that was serving these models were just fundamentally built before this and are retrofit to serve these modern models. There's going to be an entire new wave of architectures that came out and what a more exciting thing to work on than bringing this to everybody. Very different angle at the same Time how I was running a startup incubator called Prod which has incubated a bunch of different companies. Some of the earliest ones being Cursor and Anysphere which merged and Merkor and Etched went through it. A handful of others at the time as these models were getting smarter. It was 2022. I was realizing all of these companies are spending all the money they raised on compute. And I had this realization as I was working on some of my own stuff that like, oh my God, all the products I want to build are going to cost tens of millions of dollars a year in inference. This is not going to be tenable like the cost structure of every software company and the cogs is not going to be like zero anymore for an incremental user. It's going to be like quite high and it's going to be a function of inference. And then the opex of every business is going to also be inference as people use more and more coding agents. So fundamentally it seems like inference is going to be really important. And it feels like we are on a decade march for inference to become the biggest market in the world. So when you think about that 10 years from now there's going to be these giant projects where everything in that data center fundamentally hasn't been designed. Today we should go pick something and work on it. That's kind of how it got started.
Patrick O'Shaughnessy
Gavin, I'm really excited for you to go back about as far, probably early in high school, maybe even earlier, and tell your favorite hash marks on the timeline that ultimately led to your ambition to drop out of Harvard and start this company.
Gavin Uberti
My first job ever was at a company called Xnor where I did kernels development. I was 17 and a 17 year old can't sign illegally binding contracts. So rather than go ahead and do a traditional CIA, they went ahead and sat me down and said Gavin, don't share this information. And Exnor was one of the only companies that saw, hey, maybe this is a good trade and I have to work building kernels. And a number of other companies since Exnor got bought by Apple for 200 million, did the same thing at October, they got bought by Nvidia for hundreds of millions of dollars. But when you do this sort of kernels work, what you realize is that the math is relatively easy. But to get high speed decode, the thing that matters is data movement. Almost all the work that you do is optimizing. How do you move data around a single chip or across multiple chips. That's why we went ahead and built this cluster scale memory tech, we bring that interconnect time way, way lower. You can go do way more movement and as a result get a much faster time to generate each subsequent token and build these crazy things. Rob's talking about for doing a year's worth of work in a month or more than that in the future, can
Patrick O'Shaughnessy
you talk about the competitive drive that's evident in some of the high school competitions that you participated in and won?
Gavin Uberti
We did a couple. For example, I was very active in the FTC robotics. I was lucky to have a very talented partner, Safran. For a long time we were part of a traditional school team where it was about 20 guys all working together as often typical of first tech challenge. And the goal is to go out and build a robot that scores the most points and a bunch of other things too. And first they put a lot of emphasis around collaborating with other teams, around trying to do really good documentation, around trying to go ahead and get others inspired to go do the same thing. Sanford and I decided rather than go ahead and do it this way, we're going to win. And we did nothing else besides build the robot that scored the most points
Patrick O'Shaughnessy
as a two person team, as a two person team rather than a 20
Gavin Uberti
person team, that we were much, much smaller than almost every other team in the competition. We figured that if we were going to go specialize, if we were going to go out and do this, win the damn games really well, we wouldn't need to go ahead and advance based on the quality of our documentation or of our outreach. You were just going to go win. And so we did. We branched off, built this two person team that built a robot and decided we were going to go ahead and redesign it every three months. And we did. We actually had the world record for the highest score during this competition. At one point we were rated by OPR third in the world for software development. And it was a damn good machine.
Patrick O'Shaughnessy
What from that episode, can I translate as an analogy onto how you built Etched the company?
Gavin Uberti
We think about how you want to go ahead and do a full rack scale product like this. There's a couple key ideas. One of them is like velocity, velocity, velocity. That you win by shipping. You're not going to go out and win by having the best outreach or the best communications. You end up building the best product. And similarly, we think we can go do it with a lot fewer folks, that if you're willing to go ahead and just focus on product, product, product and parallelize relentlessly, you don't need 20,000 people like the big companies have. You can do the best product in the world with far fewer people.
Rob Locken
Yeah. There's a saying of the best part is no part. I think for us, it's also the best vendor is no vendor. As much as possible, we want to vertically integrate the entire product, both because we get more performance, but we can move way faster. So everything from the chips to the boards to the cold plates, to the interconnects to even the production, we want to do all of it as in house as possible. I think we're the only startup right now that's building its own rack as well as its own chips. And we did it all at the same time a couple years ago. The last time we were public at that point, we just started building our rack team and we brought over Brian Leyler, who built all of Nvidia's HGX and DGX systems, which is like 80% of their revenue. And we said, we're going to build the rack at the same time. We actually went through multiple iterations of the rack before the chips even came back. Before the chips came back, we made thermal chips that had the exact same hotspots as we expected our chips to have. So we could build the cold plates, we could over pressurize them and blow them up. We haven't had a single leak since our chips came back with the cold plates because we already validated them. We have a factory in Taiwan. We have a few dozen people out there. We built a clone of a bunch of the test stations in our office. We have a 2 megawatt data center on this floor and we did 24, 7 development cycles. People were doing day shifts and night shifts to actually get the hardware up and running as quickly as possible. It's that extreme vertical integration and extreme parallelization of the schedule that lets you get products to market way faster.
Patrick O'Shaughnessy
If you think about the building of the early team and what it required as two young guys building this company, there's lots of very talented young entrepreneurs out there, maybe for the first time of this scope or magnitude in a long time, all of whom probably could benefit from the lessons that you've learned. Getting very sophisticated, talented people to come join you even after careers at the other great companies. If you were teaching this as a class, like, here's how to get elite talent when you're young and inexperienced and naive, what would be the syllabus?
Rob Locken
We have a pretty bimodal talent philosophy. It starts with what we call the legends, which is when we're trying to solve an incredibly hard technical problem and generally do something that hasn't been done before. We need to find the very best person in the world and often the number one guy in the world versus the number 10 guy versus the number 100 guy. Huge difference in whether it's actually possible to solve the problem. We created this system we call project based recruiting where we map out all of the hardest technical problems across all industries that anyone has ever had to solve. We look at temporality. So who are the people who did the 0 to 1? Who is in charge, who actually did the work? We talk to as many people as possible and then we just track it. And you'd be surprised by the amount of people who say yes after the first conversation is pretty low. But the amount of people who say yes after the 20th conversation is surprisingly high.
Gavin Uberti
You really got to keep at them. When you hear no from somebody who really is the best in the world, then that really means, hey, should you go ahead and come back when you had a few more milestones proven out? I think it's one of the most convincing things to see is hey, we make bold claims. And when you go ahead and hit those again and again and again, that is really belief inspiring.
Rob Locken
When we decided we wanted to build a Rack and not just a chip, we were looking at this and we're saying how many products have actually shipped at scale for a rackscale system that actually have the power density that we're trying to solve? And we just kind of said if we were going to wave a magic wand, what would the best possible person in the world look like and be like, well, if we could find somebody who started at Nvidia and built the entire Rack team through all their different generations, learned all this different stuff, but is still scrappy, still understands the startup culture, but has seen scale, that would be the best possible person. So we mapped all of the different teams that related to all of the different Rack scale products Nvidia and we found three people that we thought could fit the bill and we talked to all of them and two of them have just retired. And one of them was was planning to do one more generation for Nvidia and then retire. His name's Brian and over time we convinced him to join. Brian started the HTX and DGX team at Nvidia, which was a majority of Nvidia's revenue, tens of billions of dollars a quarter. And the other two guys ended up investing, by the way. But when you have somebody like that, they just know what good looks like because they've seen it. And there's so many times where we'd talk to Brian and he'd just point to us and be like, that's a billion dollar. Like a billion dollar lesson I learned. A billion dollar lesson I learned. And that just saves us cycles. And you pair someone like Brian with somebody like Sanford.
Patrick O'Shaughnessy
Do you have a name for them? So Brian's a legend. What's Sanford?
Rob Locken
Yeah. We say chips on shoulders. Put chips in data centers.
Gavin Uberti
Yeah.
Rob Locken
So Sanford and Gavin in high school were world robotics champions. And Sanford was finishing his senior year of college. And we called him up a couple years ago and we said, hey, can you come check out what we're doing? We need some help on the platform side. He comes for a week and we say, can you build the cold plate this week? And if you asked any thermal engineer anything like that, they would think you're just totally naive. Right. I mean, these things take months to do. And to be clear, they do. But you can make real progress in a week if you put your mind to it and you think it's possible. And he built a contraption in a week that actually de risked a pretty key power question we had. And you put those two together and they've done incredible things. One is not possible without the other. Because you need the extremely driven people that just keep asking why and don't know where the bodies are buried to take tons of aggressive risks. And then you need the people who've seen scale and still have the startup scrappy mentality to help them along the way.
Patrick O'Shaughnessy
So it's really the legends, plus some naivete, raw first principles type talent. It's not just that you have both in the company, it's that they're working together.
Rob Locken
That's right.
Patrick O'Shaughnessy
If I think about that funnel, anything else more interesting to say about how much better you've gotten at recruiting and why those metrics keep getting better?
Gavin Uberti
One of the shocking things is, I'm sure being such a contrarian vet kind of self selects that if you're the kind of person who is somewhat opportunistic, trying to go join whatever the hot company is rather than go ahead and do due diligence, you will not come work here. And it's one of the things I worry about as we announce more and more of the product and its specs. We may lose some of this if we're not very careful.
Rob Locken
You kind of have to be sick in the head to join our company. Think about it. On paper, it's like you A person who is probably a very accomplished engineer making a good amount of money. It's liquid, it's predictable somewhere else you're going to convince your family to move to San Jose and live in this apartment on this housing program for the semiconductor company run by two, what, 24 year olds. Now that's pre product that is going against the biggest companies in the world and the most supply constrained environment ever created with a design that they're saying is not going to be 10% better, but it's going to be 10x better. Something must be wrong with you to do that. People are just wired differently here that they really want to not prove people wrong who don't believe, but prove people right who do believe. They just take it personally. And that's really fun to find those people. And frankly just the nature of the company makes it very easy to whittle out the people who aren't like that.
Patrick O'Shaughnessy
One of the very first things you and I talked about, Rob, was I started asking about sohu, which is the name of the first product here. And you said we can talk about that in great detail. But the thing you should know is that what we're really focused on is building a machine that can at scale produce these things and generations of them as efficiently at the highest possible quality level. So we want to build like the company or the machine that is the company is the thing that will produce this thing and then subsequent things. So I'd like to talk about a few principles or cornerstones of the company you've alluded to. Some of them you've said velocity, you've said vertical integration have become more popular topics. Parallelization is something maybe that we should talk about, but I'm especially interested in your guys willingness to take huge risk to go faster. Maybe tell your favorite story about why this is the philosophy, what it's allowed you to do that maybe other companies haven't done.
Gavin Uberti
There's a number of stories here, but one of my favorites is there was a time where we were getting close to taping out the chip. We realized wait a minute, one of our vendors is way, way behind schedule. And we had two very bad options. One option was to keep the current vendor and push our timelines out by on the order of a year. Another option was to push vendors start over and also get push timelines out by a year. Neither of these was a good option, so we had to go look for option number three. And what that was was we figured out they're all in Bangalore going and doing the work. We went out and shipped a dozen of our top engineers across the world to Bangalore for six months. I was there as well. I lived in Bangalore for four and a half months personally. And every morning we'd go ahead and walk across the crazy busy Bangalore streets into the office. We'd be the first ones in. We'd go out and built a wide variety of tools. Both things like hey, auditing a huge amount of the code that was going in, building a bunch of tools as well to make this go even faster, making sure we're making the right design decisions on the spot right there. No 12 hour back and forth, go ahead and decide immediately. And then at 1am we'd go walk back through the now empty Bangalore streets and do it all again the next day.
Rob Locken
We still had a bunch of the team in the US we ran these 12 hour on each side handoffs where we had 24 hour development cycle where at 8am and 8pm every day we'd all get on the zoom, we'd share all the data and we'd say when I wake up, this must be done. We must get this chip out. And it was extremely intense. At the same time we saw other chips at the same stage as us with that same vendor that ended up taking years that still aren't out today, they still haven't even taped out today. And it's that level of extreme urgency that's required to bring products to market.
Patrick O'Shaughnessy
What is the key to doing this? While this has become a trope because of Elon, mostly that like his special skill and others that seek to emulate him would try to do this too is figure out like what the binding constraint is and just flood the zone personally on that thing, which is kind of like going to Bangalore or something. It seems like this is a central tenet of the business and of any business that's going to do this kind of vertical integration. What's the key to doing that? Well, again, what have you learned about that specific act?
Gavin Uberti
For me, I think there are two key tricks to this. The first one is that you can't build a chip alone. It's got to be a team problem. And your most important job is to go get great people to go with you and great people to go ahead and be inspired and excited to go ahead and do crazy things like this. It is a huge ask to go say, hey guys, uproot your lives for six months or in one case 12 months that we had sent one guy out well ahead. It sucks, but we're lucky to have team Members who are in it for the right reasons. But I think the second big thing too is being able to make decisions very fast that one of the worst items is when there's a factory or there's a vendor who is waiting for you to go ahead and make some call and has been just stalled. And this happens all the time, even for very small things. Two, send folks, delegate a big amount of responsibility to them and say make a reasonable call. Okay, if you're wrong every now and then, but I would much, much rather be right most of the time and give an answer immediately than wait every time for the perfect response. Speed wins.
Patrick O'Shaughnessy
What about spending money to go faster? There's this learn by doing thing which has become so interesting. And as the world has gone away from software and towards more hardware again in the world of technology that we've outsourced so much of the learn by doing, by shipping stuff overseas and effectively just being the idea guys here in the US Seems like that obviously is reversing and you've adopted this way of learning by doing like you want to be in that iteration learning loop.
Rob Locken
Absolutely.
Patrick O'Shaughnessy
And part of that is willingness to spend and take risk with dollars. Can you talk about that a little bit?
Rob Locken
I think there's a great quote of like the biggest risk is not taking risk. Very similar here, which is like every day there's over a billion dollars in revenue in this category. And a lot of it's inference. So every day we don't ship, we're just leaving tons of opportunity on the table. So your willingness to spend money should be extremely high if you can get a very clear ROI out of it. So we have this concept that we call prefetching, which is when you're waiting for one thing to get done, when you know you're going to do other things once you have it, is there ways that you can parallelize the entire schedule? So for example, like we know our chip is going to come back on a certain date. We want it to be that everything possible that could be done without the chip is done before the chip lands. And this costs a lot of money. This means that like we want to build our entire software stack beforehand. Like we shipped racks to customer data centers without our chips in them, with all the NETWORKING, all the CPUs, all the storage all set up so we could bring all that data center software up before the chips came back. It meant that we took over 700 FPGAs and put the entire full reticle chip on an FPGA cluster. And ran a dozen different models with our full inference stack on them before the chips came back. It means that we built a thermal chip to mock the thermal profile of our chip and built cold plates based on that before the chip came back. It means we had the entire production line ready. It means we did many revs of the circuit board. It means the entire product was ready to go before the chips came back. And this is what it gives you. There is another very famous AI chip company that took 10 months to go from getting their silicon back to having them running inference in Iraq. And this was like publicly announced to their investors. And it was a really big deal. We were able to do it in 40 days. And it's because by the time the chip came back, everything was boring. The software was already written, the rack was already there, the production line was already set up. We were just go, go, go, go, go, get everything together. You don't always catch everything. You make some tweaks on the fly and then off you go.
Gavin Uberti
In that particular case too, that was a big part of it. But also the shift, I think made a big difference too. We went out and literally had a day shift and a night shift. There were team members who would come in at 10am and leave at around midnight. She would come in at midnight and leave at 10am you run it around the clock to get to that? 40 days?
Rob Locken
Yeah. I mean over half the company lives next to the office. So it makes it easier to do that type of thing.
Patrick O'Shaughnessy
You pay them to do that, right? Pay them extra. Do you still do that?
Rob Locken
The invisible hand does wonders.
Gavin Uberti
I mean, hey, it works for me too.
Rob Locken
We're both there.
Patrick O'Shaughnessy
I'd love to take one big step back and talk a bit about just the broader ecosystem here. The amount of shortages on the supply side, the exposure of risks in the global system and the supply chain around this stuff has become like everyday Wall Street Journal front page news. Like the stocks that people are watching and investing in and excited about with the memory stocks. These were, you know, boring commodity like nothing burgers five years ago. And now they're the center of global attention. If you just assess, because you've been building in it the global connected supply chain that's required to make stuff like this possible. Just riff on it, like what scares you, what's working well, what needs to change? What do you hope you change by virtue of how you build this thing? What's your assessment of this story right now?
Gavin Uberti
I think that one of the most undervalued pieces of the supply chain Story is, almost none of these things are you buy them and they don't talk to the vendor. Again, you have to go collaborate. That is the most important part, being successful. I think in ships with TSMC or with memory vendors, you need that partnership. I think that for TSMC in particular, people don't understand why it is so valuable. People look at the tech and the tech is the best in the world. But for me, the real value is all in the service. TSMC customer service is way, way better than I have seen at any other company in any other industry. It's the kind of thing where if you say, hey, you can improve your yield by making this change, you can go make them a recommendation and then we'll go run an experiment on their own dime in our case, see if they could actually get the higher yield. And when we found that we were right and the experiment worked, they moved over the rest of the line. And that kind of thing doesn't happen in most industries. If I go to like the steelworks plant and say, hey, I want you to change the composition of the steel, and they'll say, screw you, not tsmc. It is why they are the number one and why they're going to win.
Rob Locken
One of the things that matters a ton is power availability and time to power. And the problem is the more power you want, the more shortage there is. It's actually very similar to chip clusters, which is like, why is Colossus charging $12 an hour for Blackwells? It's because they're the only place you can buy 20,000 of them at once. Right. Why is the 500 megawatt data center so hard to find? It's the exact same reason. One of the things we need to think about is how do we get way more juice out of each megawatt. People are looking throughout the entire stack, whether it's just improving the pue, but also entirely new hardware to get the most tokens per megawatt to solve this problem. But fundamentally, building new buildings is hard. It's much easier to go from 100 megawatts to a gigawatt, from a gigawatt to 10 and 10 to 100. And we are pushing the limits of what's possible on these timelines. So there is a lot of people trying to scale in their data centers as much as trying to scale them out.
Patrick O'Shaughnessy
One of the interesting things about a system like this is what it replaces. Yeah, if I think about Iraq like this versus, I don't know, a set of Blackwells or something or Rubens, or whatever's coming next. How should I conceptualize that? It's not just watts, it's also physical space. The Reapers talked about this in their recent earnings call. That literally, this is the problem, that there's literally no space to put the systems. How should I conceptualize what this represents or replaces in terms of other units of compute?
Rob Locken
Here's how customers think about deploying models generally, which is when I'm building a data center or I'm building a cluster, it's not like in the abstract of like, oh, I like these chips, and this is the power footprint and so forth. It's like, I have a real production workload I'm trying to surf. And for my product to be useful, there's a certain speed I need to surf at that. And for certain products, it's really fast, and certain products it's really slow. Whatever the speed is, this is my speed. The question is, in a given amount of power, how many users can I serve while guaranteeing that speed? So another way to put it is ISO, what's called interactivity. What is my throughput? We are just finishing kind of the early innings of the AI infrastructure boom, where people really just cared about speed. GPUs were not able to reach a lot of the speeds of other types of chips, like all these SRAM chips, thousands of tokens per second. And that enabled tons of new use cases that got people very excited. There's an entirely new wave of AI chips, us being one of them, that are all going to be able to hit these speeds. The question then is, if you're hitting these speeds, what is the number of users you can serve at the same time? And by proxy, If I have 100 megawatt data center, how many software agents can I run at the same time? So when people are doing that evaluation, our hardware is going to generally be able to get you an order of magnitude more concurrency at a given level of interactivity. So that directly translates into tokens per watt, tokens per dollar, all the things people care about when they're actually serving these giant mixture of expert models at scale.
Patrick O'Shaughnessy
There's these now famous interactivity curves. Not many people publish them, but you can see a Blackwell curve, you could see an AMD curve, which is a little bit worse than Blackwell's, and It's still an $800 billion company. So if you think about what, then the impacts are of shifting that curve not just a little bit further out, but much further out.
Rob Locken
Yeah.
Patrick O'Shaughnessy
What are the things that most excite you about what this will enable.
Gavin Uberti
I mean, I want to go out and solve some of the hardest problems and I want to go solve these in much less time. There were things growing up that I was not sure I'd be able to live to see. For example, with unitistic conjecture, one of the things that we talked about in college and I was not sure I'd see that proven in my life. And this was done by an AI model and it was done over a long period of time. But if you're able to then run the same model 10 times faster, you can go shrink the time to go have these breakthroughs. And there's a huge number of other problems in math like this as well that I worry it will take 1000 years to go prove a thing like this. You can either have a much smarter model or a model of the same intelligence running much faster. You can then shrink that. And I can see it. It's so cool seeing these breakthroughs get made. I am so, so excited to see much more of this happen.
Rob Locken
I think too often people think about tasks and applications and stuff in these very short time horizons. Doing a chat and it's like 50% faster is nice, but it's not like game changing. As these agents go longer and longer time horizon and the models get more and more capable, you're going to see gigantic bodies of work that would take months of compute. And we think about this in wall clock time. If you talk to a pre training researcher at a lab, they'll tell you that wall clock time often is one of the most important things that matters. And what wall clock time means is the time from starting your run to finishing it to actually get data back. If you can shrink this time from a six month run to a two month experiment, you're going to be able to do many more iterations and people will make changes on the model architectures to actually improve the wall clock time. Very similar here in terms of how we think about the use cases. Which is the exciting part about super low latency decode is wall clock time on long horizon tasks becomes much shorter. So a year long compute build would now take a month and that month long compute build will now take three days and that three day compute build will now take seven hours and, and so forth and so forth. So that's the thing that I think is really hard to internalize because the models are just getting capable enough to do this stuff. I thought it was really cool months ago when Cursor published that they had a bunch of coding agents build an entire browser from scratch in a week. Totally nuts. And that will soon happen in under an hour. And there's going to be many of those types of things that are going to happen with these massive parallel agents all working on a given task.
Patrick O'Shaughnessy
What are the ultimate limitations of these systems? Is it just like a physics question, like how many times faster, cheaper can we get?
Gavin Uberti
Theoretically, there's a lot of room at the bottom, as they say. If you think about chip to chip latencies on an Nvidia product, you're looking at 4,000 nanoseconds to go from one chip to another. We'll be able to do much better than that. What's the mathematical limit is speed of light. You can do it in just a handful, like 2, 3 nanoseconds and they're 4,000. 4,000. Today there is a lot of room at the bottom. The same thing for things like power efficiency. That sure, we're able to go and save a huge amount by bringing the voltage down by so much, but you could go lower. You could go much, much lower. It's very challenging. But when I think about 20, 30 years in the future, then I think it's inevitable. And also for economies of scale, for clusters scale up for a long time. Eight chips was the biggest scale up domain. Nvidia even had the NVL 72, bringing it to 72. But you can be way, way bigger. You look at like a fab, for example, you have a $40 billion single monolithic building with only a handful of lines running through it. You could have the same kind of thing for some futuristic mega cluster. $40 billion, $100 billion as a giant mega token factory serving one or a handful of models for a massive number of users to get that same economies of scale thing, same model, massive number of people.
Patrick O'Shaughnessy
You mentioned kernels engineering and that being your first job that has emerged as a thing that nobody had ever heard of in their lives to now something that you hear about all the time, the importance of it, to eke more performance out of the bare raw metal. When will that just be something that AI does entirely as well? Are humans still the best kernel engineers? Are they doing it with the assistance of AI systems? Like how far down will humans still be in the loop of designing these things? Like when will that go away?
Gavin Uberti
Today it's all very hybrid and the best kernels are still written by human AI collaborations. Also, any AI models built up of these fundamental primitives, like matmules, like convolutions, like chip to chip operations collectives, and making these overlap and making these really fast matters enormously. And it's the kernel designer's job to go ahead and figure out or can I overlap? How do I allocate memory, how do I verify that if there's some issue like a retransmit, it doesn't stall the whole pipeline. These things are very challenging, but they can go make your performance be, say, 3, 4% better per optimization. You can do so many. And we thought about our software stack. We wanted to go see where the puck is going to be. And three years ago that were counted to ways you could build software. One of the most invest heavily in graph compilers. These things are not very performant, but they work out of the box. They don't require a human to go come in and tweak all the kernels. But we went the opposite direction. We are kernels first programming and that means that it for a long time did not work out of the box. But if you were a kernels expert, you could get incredibly, incredibly high performance. And the thing about this is that now as the coding models get better and better, they're doing more and more of the kernel generation task. And when the models keep getting smarter, it'll eventually do all of it. It will become superhuman. So we're going to build for where the world is going. And even today we think about our profiling tools or debugging stack, we think about it from how will the model use these tools more than we think about how will humans use these tools.
Rob Locken
We sometimes run experiments internally and we had Codex actually get GPT oss running from scratch just based off of our docs completely by itself. And it did it, I think, overnight. We think about game selection a lot and what we mean by that is making sure we're investing our energy in the right bets. Because regardless of what you choose to work on, it will take tremendous effort. And one of the things that we started with was the decision explicitly not to build an arbitrary graph compiler, not to support arbitrary Pytorch, not to support arbitrary cuda, not to support arbitrary Onyx graphs. But instead we envisioned a world where there was going to be under 100 models that actually mattered, and they were all going to look very similar from the underlying mathematical perspective. And that we were going to build primitives using physics that were going to accelerate these as much as humanly possible, and we were going to allow the most sophisticated customers to have direct access to the hardware and do whatever they want. And that has saved us a tremendous amount of time not having to build a compiler and that has allowed us to actually get much more performance. And funnily enough, when we started, a lot of people dismissed this idea. And the only people that took us seriously were in High Frequency Trading. They all hate compilers too. They all write their own kernels. And we've had dozens of people from High Frequency Trading join the team because they saw this philosophy too.
Patrick O'Shaughnessy
What are the limits to vertical integration? How do you know where to draw the line? And I'm starting with this question to talk a bit about the broader market. The circumstances of the broader market are really interesting to me. Where the vast majority of chips, of AI chips get bought by a very small set of customers, many of those customers are themselves trying to design their own AI chips. OpenAI announced jalapeno it seems like this very funny circumstance for like, the most valuable thing in the world all kind of flows through a couple chip makers, a couple chip buyers. They all seem to be thinking about doing each other's job. And then you've got the circumstance where like, okay, then these things go in a data center and you've got neoclouds and inference providers, and this other part of the stack, you've got model builders and providers. Like, I can imagine a world where because you have the best hardware, you design models and you build data centers unique outside of your current vertical. So, like, how do you think about where to draw the lines for the business?
Rob Locken
We have a saying that production is the product. Ultimately what matters here is we know inference is going to be the biggest market in the world. Whoever produces the most tokens is going to be the most valuable company in the world. So all the decisions we make is how do we get the most token capacity online as possible. And part of that is building a really good product that has way more throughput, that can run at way better latencies and so forth. So we can, per chip we make, get way more tokens online. Another part of it is like not doing parts of the stack unless we absolutely have to to get to giant scale. So there are parts that we decided to do because it was absolutely required to get to scale, like building the rack instead of just building the chips, like doing a CM model instead of doing a JDM model. But there are parts of it that are kind of noise to us right now, like we're not going and building our own data centers today. That doesn't actually help us get more capacity online in general. Our customers are actually making power and moving their clusters around to get our chips online because they're Such high throughput. If there was a world where other things were a constraint, we would totally go and integrate with them. But the reality is we're just purely focused on getting as many tokens online as possible.
Gavin Uberti
I think this comes down to economies of scale. Again, at certain parts of the stack there are huge economies of scale and others there aren't. For example, on designing models, huge economies of scale there for chip fabrication, same story. But for example, if you think about building some small metal part inside of that rack, there's not that same effect. We think the natural boundaries are on the chip side, on the bottom, as a model layer at the top, and we'll fill the whole gap between.
Rob Locken
A few weeks ago there was a guy who was running a next generation AI chip for one of the frontier companies and he's trying to recruit one of our architects and actually kind of did an UNO reverse card and started recruiting the person trying to recruit our guy. And within a week we hired him and I was going on a walk. As we were kind of finalizing the offer, I was like, you're leading this super important project. Why are you deciding to join? And his answer is super interesting, which was it fundamentally is not existential for my company for this product to win. For Google. With TPUs, the revenue comes from search.
Patrick O'Shaughnessy
Google won't fail if TPUs fail.
Rob Locken
That's right. Meta won't fail if MTIA fails, Microsoft won't fail if Maya fails. And OpenAI won't fail if Jalapeno fails. Ultimately, this is our product. It is like completely unsurprising that the best chip in the world is built by a company that only builds that chip. It's Nvidia. And for us it is completely existential for us to get as much token capacity online as possible. It recruits a set of talent and recruits a support from suppliers and from customers that view it with the level of intensity that we do.
Gavin Uberti
Look at the raw flop stats. You compare any of these chips built by the labs or by the hyperscalers. The flop density for say FBA times FB8 is lower than the Blackwell B300. And that makes sense because they don't have to go take the risk. They just have to go build a similar enough product and not pay the Nvidia tax.
Patrick O'Shaughnessy
Your finance team isn't losing money on big mistakes. It's leaking through a thousand tiny decisions. Nobody's watching. Ramp puts guardrails on spending before it happens. Real time limits. Automatic rules. Zero firefighting. Try it@ramp.com invest as your business grows Vanta scales with you, automating compliance and giving you a single source of truth for security and risk. Learn more@vanta.com invest Every investment firm is unique and generic. AI doesn't understand your process. Rogo does. It's an AI platform built specifically specifically for Wall street, connected to your data, understanding your process and producing real outputs. Check them out at Rogo AI invest the best AI and software companies from OpenAI to cursor to perplexity. Use WorkOS to become enterprise ready overnight, not in months. Visit workos.com to skip the unglamorous infrastructure work and focus on your product. Ridgeline is redefining asset management technology as a true partner, not just a software vendor. They've helped firms 5x in scale, enabling faster growth, smarter operations and a competitive edge. Visit ridgelineapps.com to see what they can unlock for your firm. As I think about you guys building the solution, the process of doing so is solving a sequence of really hard challenges. What has been the single episode that was the hardest to overcome?
Gavin Uberti
We were designing the chip. We built this massive, massive FPGA cluster to go out and verify the full chip worked as IS and FPGAs are digital entities. You can go test digital logic, but not analog logic. And it turns out that when the chip came back, we began to go see issues in our retention adaptation of incorrect results. And we realized, wait a minute, there's a problem where the back pressuring logic across a clock domain crossing is failing and this is going to cause the chip to produce wrong results and it is very, very hard to solve. We realized there was one and only one way to solve it. As we had to go line up two clock signals on our chip to within 50 picoseconds, that is literally 50 trillionths of a second. And we had to go get the signals aligned to this super small granularity and do it on every chip 2 billion times a second.
Rob Locken
A lot of people said this was impossible.
Gavin Uberti
We had people quit that people literally were like, this problem is unsolvable and best of luck guys. Well, when you have a problem like that, step one is, okay, let's assume the problem is solvable. How would it be solved? Well, first we realized what we have to be able to do is find a way to go ahead and move our clock phase by a picosecond, 10 picoseconds. And we had an idea. What if we had these two clocks just a little bit apart from each other? We figured out that hey, if you go out and Figure out the phase. If we then go out and use a drifting mechanism to go ahead and wait for just the right amount of time to get that always lined up, we can do this extremely reliably and then lock the phases exactly where they have to be. We can guarantee this never happens. People were, I think, blown away that this worked. And it worked as well as it actually did. But we made it work.
Patrick O'Shaughnessy
How long did that take?
Gavin Uberti
That was actually about two weeks.
Rob Locken
It was a dark two weeks.
Gavin Uberti
It was a very scary two weeks. But it was the kind of thing where when that kind of thing happens, that is the most important time to go ahead and be investing effort. That is the hardest time to go do it when you feel like things are hopeless. But the sooner you solve that problem, the sooner you can get back to building and scaling up production to mass volumes.
Rob Locken
I think a lot of our story is, as Gavin says, assume it is possible. Assume it is possible to have a chip with way more flops on it. Assume it is possible to have a system with way lower latency between chips. Assume it is possible to create a shared memory pool that can run at way higher bandwidth. How would one do it? A lot of the time when we do experiments, we will do dozens of experiments and all of them will fail. But we only need one to work. There was multiple times, I mean, Gavin, I think was leading the charge in our chip bring up with I think 30 different board experiments. And three of them worked. And all three of them are worth their weight in gold.
Gavin Uberti
This is one of the things people come to me and say, gavin, almost none of your experiments work. I only got to get lucky once.
Patrick O'Shaughnessy
So one idea for one of these stories that I'm asking about difficult moments in the company's history is around the ability to ra capital to fund the thing. I think when you started it, you knew you'd need capital, but you did not know you'd need the quantum of capital that you've ultimately raised and are spending to build the solution. And you hadn't raised money before. These are all new things, right? There was moments where it was really, really difficult because I was there, I saw it. Moments where it was extremely difficult to raise the money that you did. That without which the company would not exist. It would have died. And like many great stories, there are many near death moments. But money, specifically in this new world, this isn't software. You don't just need a little bit of money. Maybe you could tell the story about the true hardest part about raising money early on, before you had something that you could show people and be so proud of in performance that you could show them all their socks off. It was just you guys talking about an idea. But talk about the early difficulties raising money, because it was pretty hardcore.
Rob Locken
We've had some intense moments. It reminds me of probably early 2024, before we raised our Series A. And we were at this point where we had done enough of the architecture, done enough of the design, that we knew that the chip architecture was sound. We had to go build it. There was a lot more to do. We were ready to go into what's called the physical design stage. We needed to sign an agreement with the physical design vendor, which will cost you at least 40, $50 million. And then we had this realization, as the models were getting bigger and bigger and you're seeing these giant MOE models come out, that we were going to need to build the entire cluster, not just the chip, but we were going to need to build boards, we're going to need to build interconnects, we need to build cold plates. We're going to need to figure out all of the networking and everything, and that this was going to cost a lot more than the $15 million we had in the bank.
Gavin Uberti
And you're like, man, that was scary. You're sitting in that moment, you think, holy crap, we can't afford this. And I began looking at, like, huh? How hard is it to go back to Harvard?
Rob Locken
At the end of 23, we put together this memo. I mean, we spent 100 hours on this because we're like, we have no idea how people are going to believe us when we ask for the amount of money we're about to ask for. And it was like 30 pages. It was extremely technical and in depth of all the different things we needed to build and all the milestones we needed to hit and how the market was going to evolve and all the new use cases and the cost per token and all this modeling. And then we went and talked to investors, and every major investor in the Valley passed immediately. They were just like, okay, two kids that just finished Harvard haven't taped out a chip. No test chip inference. Who knows if this is going to be a big market? Everything's going to be training. The models still hallucinate. This could all be a bubble. At the time, the biggest semiconductor fundraisers for a Series a was around 40, 50 million dollars. We were looking at this, and we were just tallying the bill. We're like, we think we're going to spend $100 million in the next 12 months. If we really want to do this, if we want to actually get to scale and actually get the performance we're talking about, this is going to be extremely capital intensive. How the hell are we going to pull this off?
Gavin Uberti
I think that one of our key ways we got started in this process, we thought to ourselves, what is the cheapest possible way we could do this? Decided well made almost nothing. And if I ate nothing but ramen, then we would go ahead and spend basically just the money for the mass tape out and that would be that. And if so, we probably do it on $30 million, an obscenely low number. And we actually went out and got a debt provider, they'd be willing to go ahead and lend us the money we needed to go across this chip threshold. From there, I think it was first to go catalyze a series of other. Hey, maybe we can go do one more thing. One more thing?
Rob Locken
Yeah. So we're at this moment where we're like, if we really want to build this company, because we're not going to half ass it, we're not going to go do a test chip and spend years on it and, and let the entire AI market boom while we could be building the product. If we're going to do it, we're going to go all the way. We're going to need to find a way to get $100 million. I remember Gavin and I were sitting down in the office in Cupertino late at night, just looking at each other and we're like, could we cut 500k here? Could we cut 100k here? How long could we convince everyone not to take a salary? And we were like, holy fuck, the math is not going to close. We really need to solve this. And there was a period of a few weeks where you kind of just go into survival mode and you call every person that could possibly know an investor and you're like, we need $100 million to do this. If we do this, we think this could be one of the most important companies of all time. Do you know somebody that wants to take an aggressive bet, that wants to believe in us? Here's all the information. We're an open book, here's the team. It's great people. We've been working super hard. We've done these things in record time, but we have these 100 things to go, do you want to do this? And the snowball starts and you get a million here and 2 million here. And you're like, okay, we're not going to run out of money. This month, you get a $5 million check, $10 million check, and you're like, okay, maybe I can buy those FPGAs. And this snowball happened where we were very lucky that we ended up putting it all together. We had a board meeting. I show you the spreadsheet, and we look at it, and it's like 103 million. And it's like, these are all soft commits. And we all look at each other and we say, we're going to take it. And that was a Series A. And it's been much easier since then. And we've raised almost a half dozen rounds since then, many of them from those investors just doubling and tripling down. That's allowed us to get to market so quickly. Like this rack would not be possible had we not have been so aggressive.
Gavin Uberti
I also think, like, suppliers, too, I think, deserve a little bit of a commendation here that TSMC was willing to work with us back before we'd raised any of that $100 million, back when it was still really, really scary. Synopsys actually went ahead and let us get some of their emulators on extremely favorable terms, where you pay, over many years, basically a big loan. It takes a lot of belief from your partners to go do this, but at the end of that, you come out with this very strong team, and all the folks who back you are not in it, just out of pure financial incentive, they believe.
Patrick O'Shaughnessy
Why did TSMC believe, do you think?
Gavin Uberti
This is a great story. Even before you joined in, There was a conference, a semi event, and I was one of the only young CEOs of semiconductors, and I think that's kind of a novelty. Asked me to come in there and speak. I get to the semi event, and I am the only speaker there under 40 and only person there under 30. I was at the time of 22. I go up and I speak. There was a speaker's dinner afterwards, and by pure luck happened me sit next to this very senior TSMC vp. It's a very nice dinner. There's like, the former CEO of ARM there. It's very bougie. Everyone's in a suit. And I'm there with this vp, and it turns out we both studied math in college. We go out and get a little piece of paper. We begin talking in great detail about how do modern AI models work at the actual per tensor by tensor level. And the guy gets it. And we begin talking about, hey, how do you run this Very effectively. Why is level three such a critical? Technology to make this work. And the following day I get an email from TSMC saying, gavin, I want to work with Etched to find a way to make it happen. Crazy. They've been a great partner ever since.
Patrick O'Shaughnessy
It's amazing to think about some of the tropes and obviously should break the fourth wall here. I'm a big Etch Investor. I've been involved for a long time. I think the absolute world of you guys. So I'm incredibly biased in this conversation. I'm trying to ask questions that are broader and interesting and could be objections to what you're doing. And we'll keep doing that. But it's so interesting to me that when you read about investing, everyone cites this idea of contrarian and right as the quadrant that makes all the money. And it sounds really nice, but contrarian means everyone else thinks you're stupid. And so when you go and you get immediate nos from literally everybody, it is a fascinating quadrant to exist in before you become consensus.
Rob Locken
What was it like for you? I'm super curious.
Patrick O'Shaughnessy
Well, it's interesting. At the time, it was the largest by a lot. First check that I had written to say I'm not a math expert or a semiconductor expert or an AI expert really at the time. And so it was much more of a believed in the concept of this market potentially being huge. You having made very, very clear bets on how the future was going to look, having positioned the company in order to attack those things in a hardcore way. And then just the two of you and what I felt about you was the majority of the reason why we made the bet when we did in 2023 or whatever it was. But at the time it was the biggest. And I think the same thing you said about naivete applies to investing as does to maybe building a semiconductor startup, which is like, I didn't know what I didn't know. And when I called experts, they were basically like, this is stupid. They laid out in very logical terms like why this wasn't going to work and why it was such a low probability bet. And I think one of the things I've learned from it is just like you kind of have to damn the base rate. Like if you invested on base rates, you should do something other than what we and I do.
Gavin Uberti
There's always the index fund.
Patrick O'Shaughnessy
Yeah, there's always an index fund. Exactly. So it's actually never been scary for me. I think probably most of that is because there's a lot I don't know. And if I knew more about what you guys have done in the difficulty, I probably wouldn't have done it. I don't know what that says about, like, maturing as an investor. Like, maybe I don't want to know, you know, a lot more and have some of that healthy naivete. I don't know.
Rob Locken
Funny, I mean, a lot of the traditional semiconductor funds missed the entire AI chip. Like all the AI chip companies and all the coding experts missed all the coding companies. I think it's very hard to realize the constraints have changed. And when you've looked at tape outs for 20 years and you've seen so many of them not work on the first try or the second try or the third try, you couldn't even run a workload. You totally forget that EDA tools are way better and that FPGAs exist today in a way that they didn't exist before. And all the types of validation you can do today just wasn't possible for us. A lot of our believers, either they were kind of on two sides of it, they were just believers in the market and the team, or they were building chips today and extremely technical. Like the high 50 trading firms where they would literally audit everything from the micro architecture and the RTL to like the board designs and like the schedule and the software stack. And we would sit down with like 10 of their people who build their own chips. And they're asking us such detailed questions that were wondering, are they going to build the chips? It was really on either of those sides. And if you were anywhere in the middle, like, you just wouldn't understand it.
Patrick O'Shaughnessy
In the investing world, they often talk about variant perception, something that you see or believe that others don't.
Gavin Uberti
Right.
Patrick O'Shaughnessy
And that perception creates the opportunity. So I think I've invested, I don't know, five or so times in etched. And every time when you do it, stakes are getting bigger and bigger. It does get a little scarier and scarier. And because you guys have been so quiet in the marketplace, I think it's very easy to dismiss you. As the stakes get bigger and bigger, those dismissals are harder to hear. I do think betting on something that you see when what you hear from the outside world is very different, that perception gap equals opportunity.
Rob Locken
Exactly.
Patrick O'Shaughnessy
The last thing I would say is the accumulated evidence of your guys's and your team's ability to solve seemingly impossible problems is one of the most interesting things a company can have. It's like a binary. Like, companies do this or they don't.
Gavin Uberti
That's the thing. It's a big advantage of people who have Been here for a long time. You get some new joiners who are scared shitless. You see a thing like this and there are old timers who've been here
Patrick O'Shaughnessy
for all of two years smoking cigars in the trenches.
Gavin Uberti
Another one. Another one.
Rob Locken
Yeah. There's definitely a find a way mentality. If you're here, you're here because you assume it's possible. So we can't be saying it's impossible. Everything is solvable and we're just going to work at it until we figure it out. A favorite story. I have this guy who's kind of a legend in silicon validation who joined our team and we were doing the early stages of what's called wafer sort. When your chips are coming out of fab, they come out on these wafers and you have this thing called a probe card that attaches to the wafer before you dice it with these probe pads and you send these electrical signals to basically test which chips are good and bad. So when I slice the wafer into a bunch of chips, I can package it and only package the good ones. We go through our first Wafer, it's like 2, 3am because we're doing it with TSMC over the phone in Taiwan. We have the screen with the wafer that's all gray and each ship is gray. And then as you start running the patterns, the squares are supposed to turn green or red and they all turn red. We're like, fuck, like, this is really bad. Everybody's like, guys, take a breath. He leans back and he's like. The puzzle begins.
Gavin Uberti
You have to have the attitude of like, yes, you will go out and you will go stare into the abyss and you will go see scary things and we'll solve them.
Patrick O'Shaughnessy
When did you see the first green square?
Rob Locken
Within a day of that. But in the moment it's extremely scary. And there's a certain type of person who just is addicted to that feeling of just feeling the fear and solving it. And we are lucky to have a lot of those people here.
Patrick O'Shaughnessy
If you think about applying all of this earned know how from this last several years and now thinking ahead to Gen 2, Gen 3 and beyond, sure. What will you be doing most differently as a result of everything that you've learned? Just like from a conceptual standpoint, like the way that you will attack designing and producing this next one based on what you learned doing it the first time.
Rob Locken
It took us a while to get to the primitives that we think are really what matters. For scaling inference, we tried a bunch of things early on, from compilers that would turn different models into FPGAs, to burning weights in silicon, to splitting your HBM to KV cache and weights, and all of these different things. And there was a lot of cycles of learning till we got to the point that we realized that fundamentally, if you want to run a majority of tokens in the world, you need to do three things. You need to build a chip with the most flops and a given power budget. You need to build a chip that has the lowest latency between other chips, so the biggest scale up domain possible, and you need to produce as much of it as possible. And I think probably in the first half of our journey so far we learned the first two and that informed the design a lot and that informs a lot about the bets we're making in the future with the low voltage inference and the cluster scale memory. But the production part, I think in the past year has become extremely obvious how much people want to deploy this stuff. If you can have it available today, the best ability is availability. If I have 1000 chips today, someone's going to use them. And we need to build a chip that's not just way better than what's been built before, but it needs to be available at many gigawatt scale. We need to be able to be building a product that is producible at gigawatts per month in the limit. As we think about that, a lot of the design decisions we're making with our next gen, which you've seen already, is just about simplicity. Removing tons of parts, trying to assemble and disassemble a thing again and again, and learning how to make it as quick in the cycle times as possible in production, making sure it's going to be reliable, making sure it's going to be serviceable, and making sure it's going to be producible at gigantic scales.
Patrick O'Shaughnessy
What about other problems in the ecosystem that are outside of your control, such as capacity at the leading nanometer at TSMC or availability of HBM4 memory, or, you know, some of these other things where like everyone is fighting for a scarce unit of capacity or whatever. How do you face up against those realities when you're trying to produce as much as humanly possible?
Rob Locken
The people deploying the most compute in the world do think about supply a bit zero sum, which is there's only so many wafers being produced on a given nanometer node on a given fabric, right? And there's only so much memory being produced. And that's why actually for Our first gen product, we built it on a different supply chain than the Rubens. We're on 4 nanometer, Rubens on 3 nanometer, we're on a different HPM than Rubens, and so forth. So it actually is not a zero sum thing. It's a positive something where more is more. So often when we're talking to people deploying at scale, it's not a decision between a gigawatt of a GPU and a gigawatt of us, it's 2 gigawatts. And I think as much as possible, thinking about supply chain early in the design decisions, because if you have the most performant product and you can't produce it, then you're just a podcast.
Gavin Uberti
That's the other big thing about vertical integration too is. Or certain things like for the chips and for the memory, you have to go ahead and partner for most of the other stuff. Those are also very highly in demand components. And the more that you build yourself, the more stuff you can go do on top of what the world can currently build. It is not, oh, you're taking availability with somebody else. You're adding way, way more. I think that's how you win.
Patrick O'Shaughnessy
One of the things I realized we haven't talked at all about is the models themselves, which is kind of crazy. The things behind all of this demand. Anything interesting that you would say about the way that you see models progressing based on what we've seen so far? I guess I'm more interested in how you, as thinkers about hardware, think hardware might impact where the models themselves go in the future.
Gavin Uberti
One of the most important ideas that we believe in is that machines don't think like people think. You look at airplanes, for example. Airplanes don't fly like birds fly. That when you think about how mechanical devices have to work, it's often very different. And in much the same way for people, storing data and loading memory is very cheap for neurons and doing math is relatively expensive. It is the exact opposite for chips generally. 1 Data is very expensive and doing math is very cheap. And as time goes on, you'll end up finding that math gets cheaper at a rate that is faster than a memory gets cheaper. Due to this fundamental limit on any kind of DRAM device, you should go ahead and think about how can I make my model use a huge, huge amount of compute? What if I had, for example, many copies running at the same time? What if I activated a huge number of experts? What if I had gigantic experts that I could go ahead and run on multiple server acts at the Same time. That is how I think you'll build models that are the next generation of intelligence in context too. There's been a lot of work on a very efficient inference. What if I don't load the full context into memory? And most of the time I think that makes a lot of sense. You don't even build a super intelligence. Why can't it go look at a billion tokens of context? Why can't it spend a huge amount of compute to go ahead and read all at a super fast? I would love to be able to talk to a machine that was able to go attend to every book ever written. And it's short term memory. And I think we're going to get to a point where you can't.
Rob Locken
A theme in models right now is this focus on something called dynamism, which is this ability to control the level of computation and memory spent at a per token or per user level when doing attention, as well as this ability to dynamically in your chip on the fly send data to other chips for different MOE models doing certain types of operations. And the reason is fundamentally, as we are scaling context length, as we're scaling model size, as we're scaling the amount of computation per user, we're looking for ways to be more efficient. So the first thing is, like Gavin says, mixture of experts, architectures where maybe we don't need every parameter being used for every token, but maybe there's things where even at a token level we can say, well, this token needs this context. From this other token they can share that memory so we don't have to have overhead of using the memory as much. Maybe this token is really important, so we should spend more compute, we should have longer context on that token. So hardware that really accelerates these types of very dynamic computations, extremely important. And as you can imagine, current hardware that was designed before, those types of architectures have lots of overheads in doing them. So you basically end up in these really bad worlds where you have inefficient hardware at doing this dynamism. So therefore you can't run it very well, or you have these very blocky architectures that are kind of applying blunt force to many different tokens that all need more or less computation.
Patrick O'Shaughnessy
I have two questions about the future. We've talked a lot about what you've built so far and how you built it. The first is about the new ways that people might start using these systems. The raw technology, logger runtimes, things of this nature. When inference gets much cheaper, faster, more accessible there's more total supply and it's better. What are the things that you think people will use that capacity to do that are the most interesting, exciting to both of you?
Rob Locken
Yeah.
Gavin Uberti
There was a viral tweet by Noam Brown which said that as these models are having longer and longer time horizons, they can do tasks that take say six months. And there's often not enough time to go and evaluate them for such a long period of time because by that point you'll have a new model app you'll want to go evaluate instead. And with tech like we built our cluster scale memory, you can go ahead and run that six month job much faster. But there's a second piece of this too. I was talking to Noam about it. He's now an angel as well. Where it's not just the time, it's also the number of people or agents who are working on this. If you're trying to go and evaluate can a human build a rocket? You will find that the answer is no. No one person can go build a rocket. Instead you have to go put a team together. And I believe the same thing will be true of agents too. If you want to go ask can an agent go out and build some crazy futuristic piece of software? You will probably need a very large team. Maybe that's 10, maybe that's a million. You have to go and have this enormous amount of both cluster scale memory to go ahead and have that very short time per token and a huge amount of flops to be able to go run that whole fleet.
Rob Locken
I'm going to be a little futuristic. I firmly believe we are on a global march of inference becoming majority of global GDP. It may take more than 10 years, but it's going to happen. And right now we measure productivity as a society as GDP per capita, but really it's going to look much more like agents per megawatt or it may be agents per gigawatt by then. And while we're being futuristic, I think this is the second to last year where a majority of the workforce is going to be human. I think in 2027 you're going to see there's going to be more agents doing knowledge work than humans. And it's going to be extremely interesting to see what happens. You could imagine a world where for countries a majority of their energy ends up going into data centers doing inference. And the energy efficiency of those data centers basically governs how many agents and therefore how big their workforce is. So you're going to see, like as Gavin is saying, one agent or a team of five to 10 agents working on group projects for a couple days. So you can do pretty cool stuff because they're smart. But it's not going to be civilization scale. What happens when you have countries that can have literally a billion concurrent agents, like a billion people in the workforce working 24, 7 concurrently on the same stuff? It's just kind of unfathomable what's going to happen. And it's going to be the biggest proliferation of technology humanity's ever seen.
Gavin Uberti
I think as well, when you have these huge, huge amounts of demand, you get this idea of economies of scale again or think about people. I have a brain. I'm not using the whole thing all at the same time. That's only a part of it's going to be active. And this is the way healthy brains work. MOE models. It works much the same way. On the MOE model, only a small fraction of the parameters is being used for any given token at any given moment. But if you have a large number of users on a piece of hardware, you can go kind of take that brain, cut it up into many different experts on many different servers and run a huge amount of volume through it. So you'll have a bunch of different pieces of traffic, you'll have many of them using each part of the brain at a given point in time. And you'll also make the cost per thought, cost per token, way, way lower. So I think you're going to end up with these giant scale distributed brains. The form factor of this is a big data center with a bunch of chips, a huge amount of flops and a huge amount of scale up interconnect.
Patrick O'Shaughnessy
You think we'll see a trillion dollar individual data center?
Gavin Uberti
Absolutely. It is a matter of time. It's like asking will you see a billion dollar fab or $10 billion fab or $100 billion fab? It is inevitable that the economies of scale don't stop at $40 billion is the magic number for fabs. No. The cost per wafer keeps going down as you keep spending more money. And the same thing will be true of plants that go out and make steel or plants that go out and make tokens.
Patrick O'Shaughnessy
A very smart alien lands on earth and wants to know from each of you how you would frame up this opportunity that you guys are tackling. What do you say to them?
Gavin Uberti
Thinking is really valuable, that every company in the world runs on thinking. And we are entering this really unique moment in time where you have machines that can go think almost as good and as soon as good and as soon better than the best humans can. Building these machines is going to be a huge opportunity. But more important than that, the way in which you go ahead and run this kind of thinking is going to be very, very different. As demand goes higher and higher and higher and higher. There's a unique moment right now to go build a new set of solutions, a new roadmap for how do you run the future quadrillion parameter models for a billion people all at the same time on a gigantic scale up cluster.
Rob Locken
We are in a new era of intelligence where the cost of producing intelligence is dramatically so much cheaper than the value of the intelligence that we are in a many year, probably many decade supply shortage of these tokens. And basically any chip or any system that can produce tokens is likely to be extremely valuable. And you should find some part of the supply chain of the token. It can be everything from model training down to what we're doing in silicon and otherwise to spend time on and push the frontier. And that the companies that are the largest are frankly going to be the companies that produce most of the global supply of tokens and own a majority of the supply chain of that token.
Gavin Uberti
And importantly, it's people who build systems that as they get more and more chips put together, get cheaper. The way you want this to scale is not that, oh, if I want to go serve 10 times more tokens, I buy 10 times more servers. It must be some solution where if I want to go serve 10 times more tokens, then I get some economies of scale benefit with my cluster scale memory tech that allows me to then not charge as much as 10 times more for those deck tokens.
Patrick O'Shaughnessy
What a ridiculously exciting future that you guys are building to enable. When I did this with Gavin last time, I asked him my traditional closing question. So this time I'll ask you, what is the kindest thing that anyone's ever done for you?
Rob Locken
During my cancer treatment, there was a big decision I had to make. The doctors came to me and said, it's time for you to decide. Do you want to get surgery or do you want to get radiation? Here's the trade off. If you get surgery, you're more likely to live, but you have to assume you'll never be able to walk again. If you get radiation, you'll be able to walk again, but it's not the same probability that you'll live. You may die. What do you want to do? And I was 16 and my parents said, you have to make this decision for yourself. I thought a lot about it for a long time and decided, I'm going to do the surgery. I get the surgery. One of the things you do when you get a tumor resection is they do something called a necrosis analysis where they look at all the different cells and say, is a cell dead or alive? Because if you have a bunch of cancer cells that are alive, you have a problem. And they looked at it and they said, you know, you usually want 98, 99% necrosis for us to say, you're in the clear, you're below that, you should go get radiation. And there was only a few machines in the world that actually could do the type of radiation I needed. One of them was in Boston. I was in a wheelchair and I needed to move to Boston for multiple months and both of my parents decided to move out and drop everything they were doing and live with me. And I'm eternally grateful.
Patrick O'Shaughnessy
Beautiful. Thanks guys. If you enjoyed this episode, visit colossus.com you'll find every episode of this podcast, complete with hand edited transcripts. You can also subscribe to Colossus, our quarterly print, digital and private audio publication featuring in depth profiles of the founders, investors and companies that we admire most. Learn more@colossus.com subscribe.
Gavin Uberti
Foreign.
Patrick O'Shaughnessy
You know how small advantages compound over time. That's true in investing and just as true in how you run your company. Your spending system is your capital allocation strategy. Ramp makes it smarter by default. Better data, better decisions, better economics over time. See how@ramp.com invest as your business grows, Vanta scales with you, automating compliance and giving you a single source of truth for security and risk. Learn more@vanta.com invest Every investment firm is unique and generic. AI doesn't understand your process. Rogo does. It's an AI platform built specifically for Wall street, connected to your data, understanding your process and producing real outputs. Check them out at Rogo AI invest the best AI and software companies, from OpenAI to Cursor to Perplexity use work OS to become enterprise ready overnight, not in months. Visit workos.com to skip the unglamorous infrastructure work and focus on your product. Ridgeline is redefining asset management technology as a true partner, not just a software vendor. They've helped firms 5x in scale, enabling faster growth, smarter operations, and a competitive edge. Visit ridgelineapps.com to see what they can unlock for your firm.
Host: Patrick O'Shaughnessy
Guests: Gavin Uberti & Rob Locken, Founders of Etched
Date: June 30, 2026
This episode delves into the origin, strategy, and breakthrough innovations of Etched—an AI hardware company aiming to revolutionize inference by building highly specialized chips and systems. Patrick uncovers how two young founders, against the skepticism of industry veterans, brought a novel approach to AI chip design and manufacture, making inference both faster and dramatically cheaper. The conversation explores technical bets, organizational philosophy, lessons in scaling, and the macro impact of this technology for the future of AI and global productivity.
Youth and Naivete as an Asset:
"There's a certain level of naivety required to think that you could build a chip better than every other AI chip ever built..."
— Rob Locken [04:28]
Industry Doubts: Their age and lack of conventional experience led many to believe their mission was impossible.
First Principles and Breaking Old Constraints:
Winning Support:
Rob Locken
Gavin Uberti
Inference as Percentage of GDP:
"We are on a global march of inference becoming majority of global GDP...It's going to look much more like agents per megawatt..."
— Rob Locken [81:34]
Model-Hardware Co-Evolution:
Economies of Scale for Intelligence:
On Technical Innovation
On Why Inference Is So Important
On Start-up DNA and Team Building
On Being Contrarian
On Facing Supposedly Impossible Problems
| Segment | Description | Timestamp | |---------|-------------|-----------| | Intro & Motivation | The early skepticism and why young outsiders could tackle chips | [03:00–07:00] | | Technical Bets | Disaggregated inference, cluster scale memory, and low-voltage design | [08:00–15:00] | | Macro Impact | Why inference is the critical bottleneck for AI adoption | [14:34–18:30] | | Founders’ Stories | Cancer, personal inspiration, kernel optimization, and robotics | [21:00–27:00] | | Team & Culture | Recruiting industry legends + elite youth, vertical integration | [27:00–34:00] | | Speed & Parallelization | Scaling in India, prefetching, why risk is essential | [34:54–41:00] | | Supply Chain | TSMC, memory, and global power constraints | [41:49–45:34] | | Product v. Competition | Customer decision-making, concurrency, space, and power | [45:34–49:40] | | Hardest Technical Moment | 50ps clock alignment bug and break-fix mentality | [57:57–60:42] | | Fundraising Struggles | How they raised $100M+ when no one would believe | [61:38–67:44] | | Looking Forward | Building for Gen2/Gen3, mega data centers, macro AI hardware trends | [73:22–84:13] | | Alien Test | Framing the unique opportunity for a visitor from another planet | [84:23–85:45] | | Closing – Kindness | The kindest thing – family support in cancer battle | [86:12–87:33] |
Etched’s story is a case study in technical audacity, strategic focus, and the power of combining first-principles thinking with veteran experience.
Their journey overturns conventional wisdom about both people and process in semiconductors, showing that new blood and velocity can drive unprecedented advances—if paired with deep experience and relentless grit. Their vision: that scalable, affordable inference hardware will be the critical bottleneck—and unlock—for the biggest societal leap yet in the AI era.