
Loading summary
A
This is the Everyday AI show, the everyday podcast where we simplify AI and bring its power to your fingertips. Listen daily for practical advice to boost your career, business and everyday life.
B
One of the more important AI models you've probably never heard of and likely won't use just got released. This may get slept on, but Inkling is a huge release from Thinking Machines Lab. So if you haven't heard of Thinking Machines, they're led by Mira Muradi, the former CTO of OpenAI. So why is inkling maybe the most important AI model you probably won't use? Because it's now the best open source model from an American company and Thinking Machines is betting on the future of AI fine tuning as a service. The model itself, though it's multimodal, it's agentic, it's customizable, and it's available through enterprise distribution. On day one. Inkling doesn't need to beat every Chinese model or Claude fable to be relevant. It only needs to pique the interest of a few cost conscious enterprises to not only be extremely profitable, but to also help shift the conversation around enterprise AI. Frontier models are becoming so powerful they can actually fine tune smaller practical models without much iteration or without even much expertise. So has a new category reemerged? Is fine tuning back? And what might Thinking Machines first major product mean for the broader AI competition? Well, here's the big picture. Inkling could change how enterprises buy AI. So this was just released Wednesday, so hours ago, and it is not the smartest model, but it is strategically important. I think it gives companies a credible American open weight model alternative to Chinese models. And Inkling could be the best general purpose open model because it is multi modal. Right? A lot of the open source open weight Chinese models aren't. Many of them are text only. So Inkling it is multimodal, it is agentic. And as workflows in the real world lag so far behind model capabilities, I do think that there's a real market for bespoke middle of the pack AI. So on today's show, here's what you're going to learn. You're going to learn why an open model matters, even if you're never going to deploy it. You're going to understand how Inkling reopens in enterprise AI options beyond just Chinese open models. You're going to know why frontier models could make fine tuning practical beyond research teams. And you're going to know how model shopping is going to change budgets, vendors and competitive leverage. All right, let's get to it welcome to Everyday AI. If you're new here, my name is Jordan Wilson. We do this every day. It's your daily live stream, podcast and free daily newsletter helping business leaders like you and me keep up with the non stop AI updates because my gosh, can't take an hour off. I help you decide what's important. I tell you how to use all this information to grow your company and your career. So it starts here with the unedited, unscripted live stream podcast. But make sure you go to our website at your everyday AI dot com. That's your cheat code. We're going to be not just recapping the highlights from today's show, but go sign up for our free daily newsletter. We're going to be re giving you everything else that you need to know that's happening in the world of AI today because yeah, it's one of those things you gotta like go and work it out. It's a muscle. Use it every single day. And yeah, all the AI news will be in our newsletter. All right, let's get into it. Live stream audience, Good to see you. Adam joining us from St. St. Louis, Jose from Santiago, Angie joining from Montana. Amico Tokyo. Brian, what's up? Brian joining from Minnesota. So let's talk about maybe the most important AI model you'll probably never use. So there's a lot of different factors that have been compounding over the last, I would say three months. And to put it in a very short summary, frontier models are probably for the most part too much for many enterprises. Right. I'll say this. I, you know, probably the Fortune 500, they can squeeze as much juice and, you know, make it worth the cost. But I'd say for many companies, especially those probably in the Fortune 5 or like the Fortune 501 to the 5000. Right? So this isn't everyone. You know, I don't think this changes the equation for every single company, every single business out there because so many people are still going to want, you know, their chat, GPT enterprise, their, you know, Gemini, their Claude, their co pilot. They want to make things easy and they're not necessarily, you know, looking at, okay, how do open models or fine tuning models change our strategy? But for so many in our audience it does. And this is a really big deal. So here's a little bit more about the model itself. So model, the model is from Thinking Machines and it's called inkling. It is a 975 billion parameter model and it's multimodal. So text images and Audio what is interesting as well Previously Thinking Machines did demo kind of a similar dual purpose or you know, two way street audio model as well. So that can kind of hear and listen as well as talk at the same time like OpenAI's new GPT Live. So Inkling is led by former OpenAI CTO Miro Morati and its quick arrival raises I think the next big question which is how good is it? Well here's from the company themselves from their release they say Our model called Inkling is a mixture of Experts Transformer with 975 billion total parameters, 41 billion active. It supports a context window of up to 1 million tokens. It was pre trained on 45 trillion tokens of text, images, audio and video. It is the first in a family of models of different sizes. Alongside it we are sharing a preview of Inkling Small, a lighter weight model with 12 billion active parameters trained with a smaller recipe that achieves strong performance with even lower cost and latency. Inkling reasons natively over text, images and audio and balances costs with performance through efficient and controllable thinking efforts. We trained it to be a broad balanced foundation model, strong across many domains, flexible enough to adapt. Inkling is not the strongest overall model available today, open or close. Instead a combination of qualities make it a good open weights base for customization, multimodal capabilities, efficient thinking and availability on Tynker for fine tuning. Inkling is just the start, our first release in a model family we will continue to build on. We want to make customization accessible for more use cases. So Inkling is available for fine tuning on Tynker today. Picking the right base model to fine tune is a qualitative judgment that combines measurable benchmarks with a unique feel of a model that comes from playing with it. To enable the latter, we're adding the Inkling Playground in the Tynker Council, a developer facing interface for chatting with Inkling. To show what customization means in practice, we asked Inkling to fine tune itself using Tinker. The model wrote its own fine tuning job, ran it and evaluated the result. So long story kind of short, right? If you want to go use Tinkling. Tinkling, right. And I don't, I don't know why the combination of Thinking machines, Lab and Tinker, it just doesn't roll off the ton, right? But if you want to go try it out. So they do have like a dev council where you can go quote unquote chat, which is interesting that they just didn't release a chat version. Anyways, this is I think a really big deal and here's one of the reasons why. Yes, benchmarks. So podcast audience is showing the Artificial Analysis Intelligence index, but this is shaded here by our closed proprietary and our open weights models. So let me zoom way out for our very non technical audience or if you're very new to AI, what's the difference? What's proprietary? Open source, right. Proprietary are models that you can't really modify them to the core. You can add custom instructions to them and you know, change their behavior that way. But you can't really change the foundations of these proprietary models. Right. Those are the, the clauds, the GPTs, the Geminis, etc, right. Then you have these open source or open weight models and these are ones that for the most part China has been absolutely dominating on and they've been doing it through distillation, which is, you know, not exactly something the American labs are happy with, but that, you know, the Chinese labs are essentially, you know, stealing or borrowing, whatever you might call, call it the work of the American labs and making their own versions of these models and then they serve those. Right. So you can, if, if you're a big company and if you have the, the server racks, you can use these models and the open, open source, open weight models, if you have the infrastructure, you can download them and run them 24, 7 and not really pay any additional cost if you have that capacity. So that is the, the allure of these open source models or as consumer hardware becomes more capable, being able to run some of these locally. All right, so some of these models, not necessarily the ones I'm showing here on screen, but you know, some of these open source models, if you do have a very expensive, very beefy, you know, machine, you can run a slower version of them. So that's kind of the premise and the difference between proprietary and open source models. But here's why I think it's interesting because inkling where it came at and the, where it landed on the artificial intelligence index. 41. So right now your leaders are Fable 5 with a 60 and GPT 56 sold with a 59. All right, so 41, you know, seems like, okay, that's a pretty big drop off middle of the pack, right? True. But when you put into context for that, you know, go back about 7ish months, the leaders of the pack at that time, well it was GPT 5.2 with a 42. So the numbers change. All right, because the benchmarks that go into this, the artificial, Artificial Analysis Intelligence Index, those benchmarks get updated but it's essentially, you know, about a dozen or so different benchmarks that are always updated and rotated that tell you how good is a model compared to, you know, the most important factors. And Inkling, again, only being about seven months behind the frontier is actually pretty impressive for the first release from a company that we didn't really know what they were working on. Right. We really didn't start hearing from Thinking Machines Lab for their first like year that they launched. Right. So they launched, I think it was quarter one or quarter two of 20, 25. We didn't really hear anything for them for a year. And then we heard they're working on know Tinker. And then we heard that they were making their own model for fine tuning. And then we saw this, you know, this bi directional voice model. But we didn't actually see the inkling model until well less than 24 hours ago. But if you put it like that, it is about seven months behind the frontier. But it is an American model which is actually important and it's multimodal. Those two things alone I think have the potential to reshape once possible. Because my thought is right, that model right there, it's not gonna, you know, if your team is AI native, if you have, you know, agents running, if you have workflow set, you're not gonna be able to swipe like let me just be honest, right. You can't just, you know, click copy and paste and put inkling in there. But for those companies that are still finding their footing or larger enterprise companies that are looking to, you know, chunk off a big piece of their workflows to something maybe more affordable to inkling is actually not a bad option. Right. So the real product though is Tinker. It is, they are trying to turn fine tuning into a service. So inkling, yes, it can be downloaded, but even those compressed versions need 600 gigabytes of memory. So yeah, you're probably not running this or using this unless you are a large enterprise organization with your own, you know, GPU server infrastructure. So Tynker though manages the training so companies can customize the models without actually owning the GPUs. And the business model essentially turns that openness into what I think could be the first strategic reset. So let's talk about those potential strategic resets. Reset number one, American open weights could reopen procurement. So many big enterprises haven't been able to touch some of these open source Chinese models. Well, because of the current, you know, China US Relationship. Right. Especially for those companies that do business with the government Right. And we've seen the US Government get much more involved here recently between, you know, export controls. But also a huge thing here is that we've also seen reports in the past week that China may actually, which is, I don't know, funny or interesting, right, that China may shut down its models to other countries. Right. So even though they are distilling from US companies, we've seen reports they may not let you know, who knows how that will be set up, but they may not, quote, unquote, allow overseas companies to use their open source models. So I don't know how open source that actually makes them, and especially if they're just distilling them from US labs anyways, but that's beside the point. But I think for so many enterprises, they haven't been able to look at the open source category yet because of that reason. Right. If you're a company that has big government contracts, if you are using Chinese open source models, that's going to put you in a sticky situation or your RFP might be doa, Right. You may not even be considered if you are a company that has been using Chinese open source models. So that's a big unlock. All right? And I think that many companies have also just wanted to use open source models, but the whole fact that these are Chinese models and you know, we've seen reports on, you know, can you actually trust, right? Like what's kind of the, the messaging that may be coming out of this that may not be in line with companies that have stronger, you know, American values or stronger Western values? I'm not going to get into that. But you know, there's obviously different reasons, not just geopolitical reasons, that so many companies here in the US haven't really been able to, you know, convince their board or convince, you know, anyone to go down the open source route just because it is all Chinese models. Are you still running in circles trying to figure out how to actually grow your business with AI? Maybe your company has been tinkering with large language models for a year or more, but can't really get traction to find ROI on Gen AI. Hey, this is Jordan Wilson, host of this very podcast. Companies like Adobe, Microsoft and Nvidia have partnered with us because they trust our expertise in educating the masses around generative AI to get ahead. And some of the most innovative companies in the country hire us to help with their AI strategy and to train hundreds of their employees on how to use Gen AI. So whether you're looking for chat GPT training for thousands of or Just need help building your front end AI strategy? You can partner with us too. Just like some of the biggest companies in the world do. Go to your everydayai.com partner to get in contact with our team or you can just click on the partner section of our website. We'll help you stop running in those AI circles and help get your team ahead and build a straight path to ROI on gen. So here's potential reset too. The model overhang makes this customization very timely. So I've talked about this a lot over the past like three months, but I think we're now at this point where there's a model overhang. It's a little bit different than the capabilities gap that I talked about. There's the great anthropic study from. It seems like it was from so long ago, but it was only from a couple months ago their labor index report, you know, that essentially showed the capabilities of these models and then what companies were actually using them for. Right. They anonymized, I think it was 400,000 agentic chats and they mapped it all out. And essentially what they said is these models are so capable but you know, maybe, you know, enterprises are using, you know, on average about 10 to 20% of the model capabilities. So essentially these models right now, the frontier models are way more powerful than most than the average company actually needs or even has the actual capabilities to take advantage of. So that kind of adoption gap makes fine tuning middle of the pack models for stable repeated work maybe an actual new and intriguing area of AI. You know, couple that with the fact that we have seen this whiplash, the token maxing to, you know, value maxing or token efficiency whiplash where you know, earlier in, you know, from December 2025 to I would say March 2026, you know, enterprise companies were like, yes, you know, we've been all in on AI, so go use as many tokens as you can, right? And then it's like, wait, these token costs are getting higher and higher and you know, certain model providers aren't token efficient. Right. The data says that is anthropic. Right. So all of a sudden these companies have these huge API bills. So now we've seen, you know, in the May, June, July, this whiplash of companies being like, wait, we have to start reigning spend in. And part of it is, well, they're realizing that we don't need, you know, a Fable 5 type model for, you know, a hundred employees rewriting their email. And that's what I think these middle of the pack bespoke customized models could actually be a big thing. And why is that? Well, fine tuning could return because now we actually have frontier models that are smart enough to teach, right. Which we haven't really had before. I actually have, you know, a couple exciting examples that I'm going to go over here. But you know, this, the concept of fine tuning, right? So this is the teacher student model. This is essentially. Right. And we've also got reports, if you read our new newsletter every single day. We've gotten now some sniffs of real like, okay, we're at this point of recursive self improvement where, you know, models are actually according to research, right. There's actually a very interesting paper from earlier this week that yeah, the models are actually improving themselves. Right. But that also lends itself to the big, big models. I think maybe GPT 5, 6 soul is the first one I've seen in mass that we've seen actual real examples from that can teach and train smaller, much, much, much smaller versions of themselves to be used for many different purposes, right? Because these frontier models can generate data evaluations, code failure analysis. So I think it is not just, you know, tinker this fine tuning as a service from thinking machines and their new model, it's the combination of that plus this new tier of frontier models that make fine tuning a reality, right? Because it has to become common language, right? And I think we're going to see that. It's like I could right now and I have an example that I think is a really awesome one, but I think I forgot to include the photo of it in my slide. But you can if you have a decent enough machine, even a local one, right? Like a Mac studio, you can start building very, very small models yourself with hardly any idea of what the heck you're doing for very small purposes. All right, so three of these more recent examples of this concept of, you know, well, in this case, GPT5,6 Soul being able to fine tune models. So OpenAI's Jason Luai said that GPT5.6 Soul actually helped post train OpenAI's small model GPT56 Luna, completing the work estimated to require two researchers roughly two weeks. This is the one I forgot to put the screenshot in on my, on my slides here. But Pietro Sharano, hopefully I got the name right. So he's an CEO of an AI company, but he said just a fun little project that he worked on. He used GPT56 to build a small local model from his imessage history. So essentially I think he had it you know, work overnight or something like that. And it literally read every single message in his Mac history. You know, so on your Mac you can access your imessage if you're, you know, one of our green bubble friends and you're like, what does that mean? Right? So on, on my computer, I have my text messages. So he just let it go. It read everything, it understood everything, ran some models, some evaluations, and all of a sudden he had a literal small language model that GPT5.6 made for him trained on that. So this isn't like a custom GPT, right? It's like, oh, I'm using the big model to give it instructions. No, it is a separate model itself. Right. Absolutely crazy. Then Nvidia used Codex. They just put a blog post out about this, I think on Tuesday, Nvidia just used Codex and GBD56 to post train its Cosmos 3 Nano model, reportedly improving accuracy in it from 54% to 93% in one day using two prompts. All right, I'll, I'll, I'll make sure to reshare that in today's newsletter. I think we did put it in Tuesday's or yesterday's, but since I'm mentioning it on the show, I'll make sure to put it in there. But here's the reality. Two prompts. These are off the shelf. Creating or improving existing models with today's frontier technology. So it's not just inkling, right? It is more of a reflection of the combination of these very powerful frontier models that can literally build, run the evaluations, run the testing, run the qa, can literally build large, small language models on their own and also train, you know, other variations of themselves. So one example though, that we got from Thinking Machines, they did share recently one use case about Bridgewater, right? So using their Tinker service, right? So this fine tuning as a service, Bridgewater customized quin, which is a Chinese open source model. So this is, you know, obviously before Inkling was available or at least before they were ready to use it and talk about it, they did come out with this customer use case of Bridgewater customizing the Chinese open source Quinn model through Thinking Machines Tinker platform for recurring financial judgment tasks. So it report it reportedly beat the best tested Frontier model while costing 13.8 times less. So in this case, specialized judgment beat the frontier defaults. And that kind of explains this potential model shopping shift. So let me just put it out here. Y' all not want to be like, hey, I told you guys this a long time ago. Maybe I was just a little too early or a little too weird. But I think two years ago in my AI prediction series, right I talked about this very thing and the, the rise of potentially seeing thousands of small language models created by the large language models. And although we're still maybe not there because you know, there is this whole thing called, you know, compute is scarce, right? Where necessarily you can't have anthropic and OpenAI and Google and Microsoft and Meta and Grok. They can't necessarily, you know, pull gigawatts of compute to create thousands of these bespoke middle tier and small language models. But that's still the reality. And I think that this here, the combination of GPT 5, 6 soul being able to create small language models and post train its own versions and the combination of thinking machines coming up with this as a service, I think there's finally my. The I forgot if it was late 2023 or late 2024 prediction could be coming true. Where I think we are eventually and very soon going to see large language models create hundreds or thousands of versions of small language models because they're cheaper. And there's finally the appetite to care about it because nine months ago people didn't necessarily care because we were still right, we were still on this, you know, this Richie rich blank check AI, right? For $20 a month you couldn't most 90% of employees, right. If you had a team or a business plan paying 20 to 50amonth per seat, most employees couldn't get through that. So companies didn't necessarily care about being efficient with their AI. The fact, right that the concept of a small language model or a medium sized model fine tuned for specific tasks didn't necessarily matter because the big model could still do it. And everyone essentially we were playing with monopoly money until about four months ago. So the whole concept of small language models or fine tuned models literally didn't matter because it was free money. Now because of the whiplash, this is more important than ever. So as we wrap, here's what I want to talk about. Reset 3 right. Model shopping is now big and this is now I think part of the executive playbook. And we saw Microsoft, right? Microsoft is reportedly going to start moving away from OpenAI and anthropic models. Not moving away, but they're going to start mixing in their own models as well as reportedly deep seek models. So even the biggest enterprise customers are starting to realize that for a majority of day to day tasks you might not need the number one model in the world to do A big chunk of the work. And I think now it's more of this concept of maybe renting frontier intelligence for those, you know, for the ambiguous and risky and changing work. But I think the future, which I've talked about is using the right model for the right purpose at the right time and eventually that will become automatic, right? I've actually built some things like that. I'm like, I'm like, I wonder why no one else is doing this. Because it's not like, especially since GPT 5, 6 came out, like, I built skills that essentially model route, right? I can use Fable inside of Codex, I can use GPT5.6 inside of Claude desktop, right? If you have a little bit of skill and enough patience, you can do those things, right? And I've done the same things where, not that I ever hit my usage, but just a practice, right? Where I have, I've built in model routing within Codex or, you know, chat GBT work. So if you are a cost conscious, you know, company, this is the future of working with multiple models. And you're not just going to throw, you know, every single request at one big model. So here's the new AI Playbook. You own the workflow, but you're probably just going to rent the frontier or just exclusively use the Frontier for those specific purposes. So you have to score every workflow by the stakes, the volume, the privacy and how much proprietary judgment actually differentiates it. So you have to start to default to more of this economical model, right? Like you have to start mapping out that workflow and saying, okay, maybe we send the 20% to, you know, if, especially if you're using it via the API to that Frontier model, right? But I mean, my gosh, talk about middle, middle and lower tier models. I mean, OpenAI's Terra and Luna, when it comes to a cost efficiency are legit off the charts, right? So if I'm advising companies, you know, I'm saying like, you should be using these models for the most part, right? The Luna model on like max setting is probably more than enough for almost 90% of the work that you would do. So it is kind of shifting away from this one model for all purposes to understanding, which I know is confusing because it is easy to hit that easy button. But we can't just hit that easy button every single time because that button is going to start to cost more and more money as it requires more and more power, right? So you have to default to those sometimes more economical models, route by difficulty and then fine tune potentially stable, repeated measurable work. And you have to stop asking which model is the smartest and start asking which model is enough for each job. All right, that's a wrap. I hope this one was helpful. The pretty exciting release, maybe not just for the model itself, but more for what it represents. So I hope this was helpful. If so, please let me know about it. Go sign up for our free daily newsletter. Drop me a line when you get that automated, you know, email. Welcome email. But also if you could do me a favor, go subscribe to the podcast on Spotify or Apple Podcasts. Appreciate you tuning in. We'll see you back tomorrow and every day for more Everyday AI. Thanks, y'.
A
All. And that's a wrap for today's edition of Everyday AI. Thanks for joining us. If you enjoyed this episode, please subscribe and leave us a rating. It helps keep us going for a little more AI magic. Visit your everydayai.com and sign up to our daily newsletter so you don't get left behind. Go break some barriers and we'll see you next time.
Everyday AI Podcast – Ep 820: The Most Important AI Model You’ll Probably Never Use That Just Dropped
Date: July 16, 2026
Host: Jordan Wilson
In this episode, host Jordan Wilson breaks down the release of "Inkling," a new multimodal, agentic, and customizable AI model from Thinking Machines Lab, led by ex-OpenAI CTO Mira Muradi. Although most listeners may never use it directly, Jordan explains why Inkling is strategically significant for the AI landscape—especially for enterprise customers seeking American-developed, open-weights alternatives to Chinese models. The episode explores the shifting AI ecosystem, the renewed importance of fine-tuning, and how enterprises are changing their approach to model adoption and AI budgeting.
[00:17–03:32]
“Inkling could be the best general purpose open model because it is multimodal. A lot of the open source open weight Chinese models aren’t. Many of them are text only.” — Jordan Wilson [01:30]
[04:02–07:45]
“We trained it to be a broad, balanced foundation model, strong across many domains, flexible enough to adapt… Inkling is not the strongest overall model available today… Instead a combination of qualities make it a good open weights base for customization, multimodal capabilities, efficient thinking and availability on Tynker for fine-tuning.” — Jordan Wilson reading Thinking Machines Lab [07:15]
[08:00–11:50]
“So many big enterprises haven’t been able to touch some of these open source Chinese models… If you’re a company that has big government contracts, if you are using Chinese open source models, that’s going to put you in a sticky situation… your RFP might be DOA.” — Jordan Wilson [13:35]
[12:00–14:50]
[15:00–20:30]
“Fine tuning could return because now we actually have frontier models that are smart enough to teach, right? Which we haven’t really had before.” — Jordan Wilson [19:45]
[20:45–24:00]
“Two prompts. These are off the shelf. Creating or improving existing models with today's frontier technology.” — Jordan Wilson [23:51]
[24:05–25:40]
[26:00–30:49]
“You have to stop asking which model is the smartest and start asking which model is enough for each job.” — Jordan Wilson [30:30]
On the significance of an American open-weight model:
"Inkling could be the best general purpose open model because it is multimodal." [01:30]
Reading from Thinking Machines Lab release:
“We trained it to be a broad balanced foundation model, strong across many domains, flexible enough to adapt.” [07:15]
Fine-tuning is enabled by smarter “teacher” models:
“Fine tuning could return because now we actually have frontier models that are smart enough to teach.” [19:45]
How model shopping changes the landscape:
“You have to stop asking which model is the smartest and start asking which model is enough for each job.” [30:30]
For further insights and daily AI updates, Jordan encourages listeners to subscribe to the free newsletter at youreverydayai.com.