
Loading summary
A
If you showed someone a recording of this five years ago, they'd be like, we have AGI. Like it's not as smart as the models you're going to be interacting with on your computer, but you'd just be like, wait a second. This. You can have an actual real time dynamic conversation with this thing. It's insane.
B
Welcome to the Artificial Intelligence show, the podcast that helps your business grow smarter by making AI approachable and actionable. My name is Paul Raitzer. I'm the founder and CEO of SmartRx and marketing AI institute and I'm your host. Each week I'm joined joined by my co host and Smarter X Chief Content Officer Mike Kaput. As we break down all the AI news that matters and give you insights and perspectives that you can use to advance your company and your career. Join us as we accelerate AI literacy for all. Welcome to episode 225 of the Artificial Intelligence Show. I'm your host Paul Raitzer along with my co host Mike Putt. We are recording on Monday, July 13th about 9:00am Eastern Time. We had a slew of new models last week. We will get to those leading off with GPT 5.6 which sort of stole the headlines for most of the week. I don't know, it was like a day to day thing Mike, but we're going to focus on that. I don't know we're getting any new models this week. I think Google could surprise us at any moment with a Gemini 3.5 Pro. I saw this morning some possible leaks of some of the evals on 3.5 pro. So I it seems like it's fully cooked and just kind of like ready readying for release maybe. So we'll see see what the week brings us. All right marketers, get a quick gut check. When's the last time someone found you through chat, GPT or Claude instead of a traditional Google search? That shift is happening fast and most brands have no idea how they show up. Site Improve is the agentic content intelligence platform that shows marketers how their content perform across traditional SEO and AI driven search. Because good SEO and accessibility are still the foundation. They just aren't the whole picture anymore. Get your free AEO check at siteimprove.com aipod so thanks to our friends and partners at Site Improve for sponsoring today's episode. And today's episode is also brought to us by assuming you're listening on July 14th. So hopefully you you get in early, listen to podcasts early because July 14th is Macon Day. So we created Macon Day last year by asking a handful of community members, friends, team members and speakers to help spread the word about Macon. This is our annual marketing AI conference that happens in Cleveland October 13th to the 15th. And so Macon Day was designed to make the biggest registration day of the year. It was a huge success in 2025, so we're bringing it back. And in 2026, hundreds of community members, speakers, sponsors, alumni partners, friends and registered attendees are helping us make Macon Day even bigger this year. And we are incredibly grateful for everyone who's pitching in. If you've been thinking about joining us In Cleveland on October 13th, 14th and 15th, today is the best day. Again, if it's July 14th, you're listening to this, the best day to register, you'll save $200 on your pass or with code Macon Day 26. That's M A I C O N Day 26. And you'll be entered for a chance to win exclusive prizes including a VIP party pass, a VIP lounge pass, which was a huge hit last year, a complimentary on Demand upgrade, and $100 in Macon swag bucks. I didn't even know we had Macon Swag bucks, so that's cool. I might. I might enter to win. I'm going to register today. The offer ends at midnight Eastern time, so don't wait. Visit ma day.com so that's M A I C O N day.com if you need more information on Macon before committing, a link on the Macon Day website will take you to the event site. So thanks for being a part of this community. We hope to see everyone at Macon this October in Cleveland. Okay, every week we start off with a recap of our AI Pulse survey. This is an informal poll of our listeners on how they feel about topics that we covered in the previous episode. So last week we had Palantir CEO claims AI Labs quietly absorb your company's data and competitive edge. Do you worry about this when using AI providers? 48 somewhat. But we accept the trade off. 28 yes. It's a serious concern for us. 11% no contracts and controls protect us. That's. That's kind of what we've always thought. I'm not so sure about that anymore. And 13% haven't thought about it. All right. And then the second one was, what is AI actually doing to headcount at your company right now? 59% no noticeable impact on headcount yet. So all the economic reports would support that as the majority right now, 30% were hiring less or slowering hiring because of AI. We are definitely seeing and hearing that 9% were hiring more because of AI. That's great. Those are probably our high growth companies that we talked about last week. So you're growing and you're hiring as a result and then a small percentage. Not sure. All right, so Mike, we have, like I said in the opening, lots of model news to discuss this week. We're going to focus on what's new from OpenAI. And OpenAI had a very busy week last week, but we're going to start with the models.
A
Yes, they did, Paul. So first up, OpenAI had some of these big releases this week. So first we saw the general launch of GPT 5.6. This is the company's most powerful model family yet. So this comes in three tiers. There's GPT 5.6 Soul, which is the flagship, built for complex work like coding, research, science, computer use, GPT 5.6 tera, this is a middle tier, balancing capability, speed and cost for everyday work. And GPT 5.6 Luna, which is the fastest and cheapest of the family. So all three of these are now available generally across Chat GPT, Codex, the Open AI API. And they are positioning this model family as the new frontier. The company says Seoul itself sets a new state of the art on the Artificial Analysis Coding Agent Index. It gets a score of 80, which they claim is 2.8 points above Anthropics Claude Fable 5, while using less than half the output tokens and costing about a third less. Now, this launch itself is noteworthy not just because of the model, but as we've discussed. The Trump administration pushed Open AI into a staggered release last month. They limited initial acts access to these models to government approved entities, while the Commerce Department's center for AI Standards and Innovation tested the model. These restrictions were lifted this past week, though a White House official disputed that any approval was needed, saying decisions on timing and scope of releases rest entirely with the companies. Now, not only did they release these new models, but OpenAI also launched something called Chat GPT Work. Now, Chat GPT Work is an agent that takes an outcome, gathers information across your apps, and stays with a complex project for some duration of time. The minutes could be hours, OpenAI claims, breaking it basically into steps and completing all this on its own. So the output that this tool produces is finished work sheets, slides, docs, shareable web apps, etc. It can it can connect to the systems where work already lives, including Slack, Microsoft Teams, Google Drive, etc. This launch also reorganizes OpenAI's product lineup and we'll kind of talk about this a bit, but the Codex app is basically merging into a single new chat GPT desktop app that puts chat work and codecs on every plan, including free. OpenAI is also beginning to sunset its Atlas browser, which is its kind of agentic browser it was experimenting with. And on top of all this, there was a new wave of voice models. They released GPT Live, GPT Live 1 and GPT Live 11 mini. These are models that are built to listen and speak at the same time when you use Chat GPT voice mode, so that conversations flow naturally and users can interrupt without the model cutting off. Interestingly, open AI's chat GPT voice productly told Axios the company thinks this will unlock the ability to use voice as kind of the primary interface to computing. Now, if that wasn't enough, the rumor mill is already spinning dramatically. There are posts circulating on X partially corroborated by AI commentator Andrew Curran, who we've talked about that claim. GPT 5.6 is the final model in the 5 Series, and that GPT 6, built on a much larger base model, could actually arrive within weeks. So, Paul, there's a ton to unpack here. Maybe kick us off with your thoughts on the initial releases models. I know there's probably some stuff to talk about chat GPT work as well.
B
Yeah. So I've played around a little bit with 5.6 over the weekend. I mean, I haven't done a ton like pushing it on, you know, internal evals or anything like that, but definitely seems like a significant upgrade over 5.5. Early response has been really strong. People who had early access, you know, seem to be really happy with it. A lot of people seem to prefer Fable 5 when we're, you know, if you're comparing those as direct model comparisons. So according to OpenAI 5.6, Sol sets a new standard for intelligence and efficiency, achieving state of the art results across coding knowledge, work, cyber security and science. They say fewer tokens and at a lower estimated cost, but everything I've been seeing online is that the thing burns through tokens like crazy.
A
It can, yeah.
B
Like people are running up against the limits like really fast. So that's just something to keep in mind. I did see some things this morning that people are complaining that Open the Eyes sort of like nerfed it. Since it first came out like four days into the launch, it's already seen a performance decline and I even saw some comments from people within OpenAI that they're basically working on how it functions behind the scenes. And it might cause some issues with its reasoning capabilities, stuff like that. The early access from governments is super confusing, as you alluded to. Ashley Gold at Axios had the story that OpenAI got the green light from the government and then posted later that day on July 8th that the White House official disputed that they gave a green light and that they don't need that permission, and referred axios to the June 2 executive order, which bars any mandatory federal licensing or pre clearance. So it's like, it seems like maybe there's just a double standard that if it's anthropic they have to get clearance. I don't want to be like overly negative or suspicious here, but like maybe offering 5% equity in your company helps when it comes to getting quick approvals of model I don't know. So who knows what's actually going on there? There was a Matt Schumer who we've talked about before of the Something Big is Happening fame from February when he posted that article about like the impact of coding models. He tweeted that 5.6 SOL just accidentally deleted almost all of my Mac files and this is why I trust Fable 1000 times more. He then said, the crazy thing is if you read my GP4GPT 5.6 sole review, I already much preferred Fable and stopped using 5.6 weeks ago. The only reason I was using it today is because OpenAI team asked me to test the Ultra mode. For what it's worth, they're great to work with and it's a freak accident, just sucks so much. So again, Schumer obviously is pretty advanced user of these models and even he, for whatever reason, almost accidentally had his entire Mac files deleted by a model. So user beware. One of the exciting things about 5.6 soul is it renewed the Sam versus Elon Twitter battle. So that's always fun to see. So Sam tweeted, there are a lot of benchmarks that suggest 5.6 SOL is the best model in the world right now, but the most reliable way to tell is that Elon is obsessed with me. Again, this was in a reply to something we'll touch on a little bit later about Apple suing OpenAI. But Elon had tweeted he takes scamming to a whole new level. To which Sam retweet or replied, homeboy, you're the one selling public market investors on short term space data centers. To which Elon replied, we start flying them next year. Maybe you can come see them after your parole officer approves after stealing an open Source Charity. You then stole all of Apple's phone technology. What do you plan to do for an encore that's tough to beat? So, you know, we haven't had the soap opera of Sam vs Elon for at least like six weeks, so. So it's always good to have that back. Okay, so chatgpt work, honestly, and Mike, maybe you can break this down for me a little bit more. I, I find this really confusing as to when you're supposed to use which thing. Now. So if, like, if we go into Smarter X's ChatGPT instance, you can just click whether you want chat or work. So I'll just kind of walk through a little bit of what this is like, if you haven't seen this yet. So in our account, you know, if I go in and do a normal chat. So if you're what, you're a Claude, user, Gemini copilot, whatever, you have your chat window and at the top there's a chat and a work tab. So if I'm in the chat tab and I'm going to start a conversation in chat, my model dropdowns is instant 5.5. Then there's medium high or pro of instant 5.5 or, or I guess 5.5. So you have instant medium high pro of 5.5. Then I can choose 5.6 SOL. Underneath that I can still choose 5.4, which it says leaving July 23rd, 5.3, and I can still select O3. So in a traditional chat, those are my options. When I click into the Work tab, my model choices are now 5.6 SOL, Terra, Luna, or 5.5. I can also then choose Effort, Light, Medium, High, Extra High, or Max. And I can choose Speed, Standard or Fast. I can then choose a project, so I can connect work to a project. I can also connect plugins. Now one of the parts that confusing to me is you could already do plugins with chat, so it just seems like it's the same feature. You could also connect projects with chat. So, like, these don't seem like differentiating features. And then there's a call to action to get the desktop app. And I'm still not actually 100% sure what the difference is between the desktop app and using the browser, but that's a standard software thing. Like I always use the browser for Asana as an example, not the desktop app. Okay, so that, that's kind of like that. I'm going to now I'm going to give the prompting and then maybe Mike, I'll stop and ask if you have any clarity on this that I'm missing? Okay, so then OpenAI has a prompting guide to try and differentiate how to prompt when you're using work versus Chat. So in chat it says a short prompt is often enough for larger or more important tasks. Include the parts that matter, like the goal. What should ChatGPT do? Context what information or sources will help output? So what format, length or level of detail do you need? And then boundaries what must stay unchanged? What should ChatGPT avoid? So that's how it guides you to to to work with traditional chat. Then it says for prompting work, use chat for quick questions, short rewrites, brainstorming and lightweight drafts. Use work for tasks that draw on different sources or tools, involve a sequence of steps, make changes, or produce a large deliverable for work tasks, describe the results you need, provide the source material, name the audience, and explain how you review the work. Ask ChatGPT to plan, gather the needed information, create files and check them before it finishes. Work is useful for time consuming or recurring tasks or for finished files you can reuse. A task that uses more credits can still be worthwhile if it saves time, improves quality, or helps you make an important decision. Start with one result you can review. So include the relevant resources to find the audience, separate required work, blah blah blah. Review the first result, refine the instructions and reuse the the workflow when it works. And then one other note here Mike it says ChatGPT work is an agent that takes an outcome, gathers information across your apps and stays with a complex project for hours, breaking it into steps and completing them on its own. This can be a little misleading because projects that takes hours, there's, there's no way that the reliability is high on that stuff, right? And it's going to burn your entire token budget. Like if you're using work or an agent to do hours of work it it's going to burn through whatever your credits or token but whatever it is so then they said in the launch ChatGPT work is designed for longer more involved work than a typical chat request. So usage works differently. Usage varies with the amount of work required, and more complex tasks may use more of your plans included Usage ChatGPT follows the same usage structure as Codex, Chat GPT Enterprise and EDU Admins can set spend controls in the admin console to manage chat GPT work usage as adoption grows. So I'm just going to stop there Mike, honestly like so as an account admin for our Chat GPT instance and as a user who regularly, you know, a Dozen times a day is in Chat gbt. Using it for different things, I have no idea because most of what it says for work, I do that with chat GPT standard. So I'm not 100% clear when I'm supposed to jump over to work or if I'm just supposed to now stay in work. Do you have any clarity on this?
A
That's shocking to me, Paul, because in true AI lab fashion, they've made this as confusing as possible, I would argue. Here's what I've observed. I don't have a full read on this yet, but just some tests I've ran. I think the two key distinctions are first, the ability to spin up sub agents, which for instance, I just tested this in the web app. I said, hey, can you go like research current AI capabilities for me, spin up some agents to do it? So it spun up some sub agents to apparently do this task and worked for a bit. I don't think you can do that in chat. That can be helpful for parallel work. Here's the bigger thing though, which I just tested and is deeply confusing to me. In the Chat GPT desktop app, this thing has the option to use your computer and that is the key differentiator. That's what Codex was able to do so well is that now once I enable computer use within this app, this thing can now go do all sorts of stuff with my files, spin up agents to do all sorts of things
B
that are delete them.
A
Very, very dangerous. Do not enable this if you don't understand the capabilities or if you're not allowed. If you're not allowed. Correct. So that is a key differentiator. However, the web app, as far as I can tell, does not have that ability. So I was like, for me, I was using the desktop app, just looking at this and I was like, oh, okay, this kind of makes sense to me. This is just like Codex Lite for non technical workers. It's like Claude Cowork, basically, I think is like the analogy here. However, that only holds true in the actual desktop app. In the web app, sure, it seems more agentic, but it doesn't seem to be able to use your computer or anything, which is good in a lot of cases, but doesn't differentiate it as much, I would say. So I'm a little confused if they're just trying to drive you to the desktop app because you would think long, long term computer use is like the name of the game so that they can then do knowledge work for you. So that's kind of how I've looked at it, but that's why I was very surprised. This seems like a monumental change and I don't know if people are treating it that way because now, right now in your chat GPT app, you have the ability to have this thing, give it, give this thing access to your computer. So like if you're an enterprise that doesn't automatically have that shutdown by it, like you need to go in and figure this out Now. I don't even know if you can restrict it. I have no idea. I'd have to look in our account. But that seems huge to me and I just don't know why the labs might wouldn't say.
B
It's like their go to market plan doesn't match the significance of the launch. It's like, hey, we launched this thing, here's a couple blog posts. And then you go and it's like, wait a second. This is like entirely different ways of talking to these things, right? And the one I keep coming back to is like when I reread this, it's like use chat for quick questions, sort rewrites, brainstorming and lightweight drafts. Use work for, you know, a bunch of things in larger delivery. So I'm thinking the majority of my use of these models is strategic support and planning and like using the reasoning. It's like, is that should I be in work instead of traditional chat? Like I don't, I don't even know. But based on their explanation here, the answer would be yes, that chat is literally just for like the really quick stuff. Quick stuff and work is where you live if you're using reasoning or agents.
A
In essence, that's my understanding. Yes, you're still using the same base model, so I think you could still accomplish a lot of the same things in chat is my guess. But again, I don't know for sure.
B
All right, well then to make things more confusing, Ethan Malik, who had early access to both ChatGPT work and 5.6, he tweets. I've been going on about how ChatGPT work and Claude co work are missed opportunities for knowledge workers. And to illustrate that, take a look at Google's Notebook LM answering the same question as ChatGPT work with the same 70 files. It centers process and sources, not just outputs. To be clear, Notebook LM has its own issues and is built for specific use case research and analysis of sources. But as an example of how a UX might actually operate, that treats knowledge work seriously. It doesn't treat only goal as outputs. It exposes processes. Now the point he was making was he showed a screenshot of an output, I think it was from chat GPT work where it was like here's your file, like here's your PowerPoint and where Notebook LM has this like extensive dashboard of all these different capabilities and click and look for references. So then he was sharing on top of a post he had previously put up and I thought this was helpful. He said a fundamental problem with extending Codex cowork code to all knowledge work is that they remain very software brained where the end result, the software is what is important and that code serves as the source of truth. For a lot of other knowledge work, the process is at least as important as the outcome. This includes researching what is known and exploration of alternatives, failed efforts, prototype branches, experiments, etc. All of those things are valuable. So you cannot use the PowerPoint at the end the way you can use a code base. Nor is progress on a to do list sufficient context. Post compaction you work in learning loops, refining your perspectives as you go. In some ways this makes long running models like Fable hard to use for deep knowledge work since they are designed to deliver product to you. At the end you can prompt your way around this problem. But everything about the codecs and code harnesses want you to be a software developer and you have to fight them. I think that's a really, really important thing to keep in mind is like cowork and work are being powered by the coding agents underneath them and the harnesses that structure those which are built for software developers and AI researchers and they're like force fitting them to the rest of the world now all these knowledge workers, but they're still being built by software developers who understand software. So he said there's a real disconnect between how a manager or analyst thinks about problems and how the Agentix software tools approach solving them. Addressing this is critical to breaking out of the coding niche for these tools. So I don't know if you have any other thoughts there Mike, but I just thought that was really important context from Malik on maybe why it isn't super clean how to use these things and when I couldn't agree more, I
A
would say the success and value I've gotten from those tools is just prompting around these behaviors or doing stuff that lends themselves to those behaviors. So yeah, it's deeply confusing and I also just come back to both the challenge and the opportunity. So many non technical knowledge workers don't understand just how powerful these tools can be for specific types of work. But like, like, who's going to tell them? How are you going to communicate? There's an actual sea change here moving from chat to computer use agents that's deeply important for people to grasp. But you wouldn't know it from any of these announcements.
B
Yeah, I will tell you like just some inside information how we think at SmartRx. So our AI Academy consists of dozens of, you know, on demand professional certificate courses that are, you know, four, five, six courses deep. They might take three to five hours to complete and you earn your certificate on the other end. And we see that being continuously incredibly valuable within organizations. But the dynamic nature of how fast these things are changing and like, what it means to all of us as knowledge workers. Had us sitting there Friday morning literally discussing our roadmap for AI Academy and like, okay, how do we address the fact that these things are changing so fast? And we have our weekly app reviews that we drop app and agent reviews that come out every Friday. We have lives happening all the time. But like, there's a whole nother velocity happening right now behind how these models work and how it changes the way we work. So. So we are very actively thinking about how to continually evolve what we're doing with Academy to address the fact that there needs to be more real time learning that on demand course and certificates are not going to be sufficient on their own that you really, it becomes what, what changed last week and what are we going to learn from it? So we have some really cool things in the works, but I feel like every day it's becoming more and more urgent. And as I was preparing even for today's episode, I was just like, oh my gosh, I want to go build the next iteration of what we're going to do right now. Right.
A
I couldn't agree more.
B
And then a couple quick thoughts on GPT Live, because I do think it's, it's huge. It's, you know, an indication of what OpenAI believes that voice is going to be the future. They said our vision is to enable truly natural human AI interaction. A world where collaborating with AI feels as fluid and responsive as working with another person. While reasoning and complex task execution happened seamlessly in the background, the way they're achieving this is kind of interesting. They had a post that talked about how their previous voice models worked and there was two prior generations in essence, and a lot of voice models work this way. This isn't just open AIs. So Cascade voice systems, which is, you know, a previous generation, is kind of how Siri works. Rely on a series of models acting one after another to process each turn. So original chatgpt voice chained three models together. So there was a speech to tie text model. So you would talk to the model, it would convert it into text, so it transcribe what you said so it could then understand it. Then a large language model would produce a response in text and then that there was a text to speech model that would convert that back to speech. So if you wondered why talking to Siri or other models is slow, it's because there's a symphony of things happening behind the scenes to power that communication. Then there was turn based voice models like ChatGPT advanced voice mode, which processed and generated audio with a single model. So that was the breakthrough we talked about last year. This reduced latency and made conversations smoother, but it still operated in discrete turns. That's why like you'd be talking and you take a breath and it starts talking back to you. It's like, oh, hold on, I'm not, I'm not done telling you what I was going to tell you. So they say GPT Live addresses these limitations through two changes. Instead of processing a sequence of separate messages, Live continuously, continuously processes input while generating output. So it's listening while talking. In essence, the model can therefore make interactive decisions many times per second whether to speak, continue listening, pause, interrupt, or invoke a tool. And then second, they decoupled Live, which handles continuous interaction from deeper work. So when a question requires search, reasoning or more organic capabilities, Live delegates that task to another model like 5.5 right now, and eventually 5.6. This allows the conversation to keep going while it's doing tasks in the background. So as a result, conversations should start to feel much more natural. You'll be able to interrupt with a question pause to go through your thoughts. And it shouldn't feel as sequenced. I guess it should be happening kind of simultaneously. So should be really interesting. They say that these models are rolling out as of last week to ChatGPT users globally. So if you haven't tried voice recently, might be worth it, you know, pop in and do that. I know, Mike, you're a huge voice user, but I know you use Whisper all the time for transcription, but I assume you also are talking to the models a lot too.
A
Yeah, I've used this quite a bit. I so far really enjoy it. It's got its flaws, but it is pretty night and day from the previous voice mode, which even the previous voice mode I still found really valuable despite its limitations. But yeah, this one, it's like if you showed someone a recording of this five years ago, they'd be like, we have AGI. Like it's not as smart the models you're going to be interacting with on your computer, but you'd just be like, wait a second, this, you can have an actual real time dynamic conversation with this thing. It's insane. I, I really hope they crack the code on like these models being smarter and able to use tools, et cetera. I realize there's like some technological bottlenecks at the moment, but like the moment you're able to just talk to these things and say, go code me this, go access this tool, go do this, that and the other or whatever. I think productivity goes crazy if you, if you are someone that tends to use voice a lot.
B
Yep. Yeah. And if we'll put the link in the show notes to their post. It does. There's actually. You can listen to comparisons of the previous generation, the new generation with some sample voice things. So yeah, it's, I think it's a, it's definitely the start of a new generation. My, my assumption is Gemini is or will soon function the same way because generally they've been been together in the lead here. But Google is certainly very advanced in terms of voice because they've been integrating it into search and other elements of their business. So yeah, definitely an area to keep an eye on.
A
All right, our next big topic this week, highly related. We're talking about some advice and considerations around both agents and just AI generally in the enterprise. So this kind of kicked off with this past week Box CEO Aaron Levy. We've talked about a bunch. He published this widely shared rundown of what he's hearing from enterprise IT IT leaders about AI agents coming off a bunch of meetings he's had with several people in these roles. So he basically gives this kind of map of the real unglamorous challenges that companies hit as they're trying to use agents in production. So Levy's biggest theme here is that agents basically force this like operating model problem. So most companies are built in silos, but agents work best when tied to a process and the most valuable process is cut across these silos. So this raises a bunch of like very difficult questions for enterprises to answer, like who owns and manages centrally deployed agents and how do they actually get adopted across organizational boundaries. So he talks about some things like data fragmentation, being a major blocker underneath all that, since agents struggle to give accurate on policy answers when a company's data is scattered and non Standardized. He argues that in a world where everyone can tap into roughly the same frontier intelligence, a company's proprietary context, which is its own data captured and formatted so agents can use it, becomes its real competitive moat. He also said there's a growing consensus that tokens are the wrong metric. Companies should manage instead to business outcomes like revenue or shipped product. The catch being those are much harder to track top down. He also added two more themes that are important here. The best use cases fundamentally change the work being done, rather than just doing an old process more efficiently. And the talent to deploy and manage agents agents is very, very scarce right now. Most companies, he says, will have to train for it internally. Now, on top of all this, we got BCG's fourth annual AI at Work survey of nearly 12, 000 employees that found the AI is changing jobs faster than companies are redesigning how they operate. So they found 74 of frontline workers are now regular AI users. That's up 23 points from last year. 61% believe agents could do at least half their job within three years. But overall, their baseline basic finding is that strategic clarity, not access to tools, is what separates the organizations Getting real value. Paul that was like news or music to my ears because that is the exact approach of the AI for Productivity workshop that we're doing at the boot camp this week, which we'll talk about more in a sec. It's just like the tools of course matter and literacy with the tools matters. But all these enterprises, it seems, are running into all these bigger challenges about workflow, mapping, context, governance, et cetera. Like, what's your advice right now for enterprises? Like do you see these same things in the conversations you're having?
B
Yeah, this is a really good report. There's they surveyed 12,000 frontline employees, managers and leaders and dozen global markets. So it's actually like a lot of international components as well. A lot of this mike reinforces themes we saw in our state of AI for business research that we released in was that May. I think we came out with that search research. So yeah, I'll call out a few points here. So they said 42% of AI users save eight hours, the equivalent of days worth of work in a week. And the time savings is even higher for functions such as marketing 60% IT 53% and human resources, 50%. However, 66% still receive little or no guidance on what to do at the time they save. And more than half say they're not reinvesting time saved in more strategic work. That goes to the point, Mike, you were making about just the organizational structure of it's like, okay, great, we gave them tools and maybe we even trained them how to use the tools, but we didn't train them what to do with the time they're saving. And so they're not redirecting that into strategic efforts and new campaigns and new ideas and things like that, which is a huge opportunity for companies. Compared to 2025, number of organizations that have graduated to using AI to reshape workflows end to end and to invent new business models has nearly doubled from 40 to 42%, up from 22% you mentioned the idea of training and support from leaders is a strong driver of AI's potential, but also one of the biggest unmet needs. 72% of respondents say expectations for the skills they need have shifted. However, 36% feel they have received, they have not received adequate or they have received adequate upskilling. So like a smaller percentage are getting the upskilling they needed. And then only a third of frontline employees say leadership's communications about AI are clear. This is what we see all the time. And only 28% see a strong connection between what leaders say and what the organization actually does. So we talked about that when I did the eight pillars of business AI transformation. I said like the most important thing was clarity from leaders, that leaders understood the moment and that they were clear in their communication. So definitely echoes the things we've been saying. And then on the topic of agents, the vast majority of respondents, 84%, have heard of AI agents tools and that act autonomously with minimal human oversight Site is how they define them. More than twice as many respondents as last year say their organizations have integrated AI agents into workflows. So that's at 30%, up from 13. Another 50% say their workspace has run AI agent experiments or pilots. The experience is leading people at all levels to believe that agents could do at least half of their job within three years, with leaders and managers expecting the biggest shift. That's. That's crazy. Yeah, I mean, that's. They're right. But the fact that people are now realizing that awareness and integration of AI agents have outpaced the systems that companies have enacted to supervise them. This is an issue that Levy talks about. Half of respondents say their companies lack clear governance for managing teams with people in AI. And almost as many say AI related accountability is one of their three top concerns of the future. And then they had a really nice five CEO imperatives outline. So I'll say The five points. So make strategic clarity a top priority and own it. Personally I agree 100%. The CEO has to be the lead on this stuff. Change the scorecard. Measure value, not adoption. Invest in redesigning work end to end, not in more tools. Put people at the heart of that redesign and then govern it as a moving target, not a one off program. So those are all really good points from for Levy's overview. Couple other things I'll drill into and and maybe double click on Mike. So growing view that enterprises are going to live in a multimodal multimodal world world. Lots of interest though early in the actual adoption in layers that can route workloads to different models for cost performance. So this means like you know, let's say Claude goes down. What are we going to do if all of our everything lives in there, all of our projects, our skills, whatever, how do we keep the business running? Or if it becomes too expensive, we need to have another. We need to have an open source model internally. You need to have a secondary backup. So he's basically saying like almost no enterprise is going to bet on a single company or model provider. You mentioned this one talent for driving AI adoption and implementation still remains a major issue and topic. Many view it as something you necessarily have to train for internally do a shortage of talent being trained on this from the outside. So this goes back to our whole point like AI literacy is fundamental. Like the the companies have to own, reskilling and upskilling their own people. It's going to be the fastest path, people with domain knowledge or institutional knowledge, domain expertise that you can make AI literate, AI forward. That is your best path to do this. Right? Right. And then the best use cases for AI tend to be those that fundamentally change the work being done instead of just replacing an existing process. So this goes to the idea that I've been touting that I've, you know, featured in my Macon 2025 keynote. Optimization is doing things better, faster, cheaper. That's 10% thinking. Should we should absolutely be doing it. But innovation is reimagining what's possible and creating entirely new forms of value. That's 10x thinking and that's what we want to be doing now. One last late entrant to the game, Mike. This was from Satya Nadella on Sunday, July 12th. So last night he posted this. So I just threw this in and kind of the last minute he published something on X called the Reverse Information Paradox. It had seven and a half million views as of Monday morning. So this goes along with what Karp was saying last week. You know, Palantir CEO, that we talked about a little bit of what Levy was talking about, but you're seeing these recurring themes. So he said, in the age of intelligence, how should firms protect their core ip? Nobel Prize winning economist Kenneth Arrow famously described a paradox in the market for information. Quote, its value for the purchaser is not known until he has the information, but then he has in effect acquired it without cost. So in Arrow's information paradox, the seller risks giving away knowledge in order to sell it. AI creates the reverse problem. In the AI age, the buyer risks giving away knowledge just in order to use what they bought. Meaning you're giving these model companies knowledge every time you use their product. So he said you essentially pay for intelligence twice, once with money and again with something even more valuable, the proprietary knowledge you must reveal to make that intelligence useful. The better you want the model to perform, the more of that knowledge you have to feed it. The seller learns, meaning OpenAI, anthropic, Google, and in theory, Microsoft. The seller learns more and more about you as you use what you purchased, while you learn very little about what the seller is learning in return. So you don't know what your prompts are teaching them, what the outputs you create are teaching them is the point he's making. That is what I think of as the reverse information paradox. Models learn from exhaust the prompts people write, the tools agents use, and especially the corrections people make when the model is wrong. Every correction is distilled into institutional know how in consuming intelligence, you are creating intelligence and what you create should belong to you. This is a very different approach than like, I mean, Microsoft is the biggest investor in OpenAI and they're in essence taking the anti OpenAI position here, which is really fascinating to watch happen in real time in learning flows. If learning flows in only one direction, economic value converges towards the owners of the learning infrastructure rather than the creators of the knowledge itself. So again, the model companies win, not you. Therefore, it's imperative that we distribute the learning infrastructure to every firm so that they can control their own learning loop, as Alec Karp put it. Again referring to the Palantir CEO topic we talked about last week. What the technical customers want is control over their compute, their models, their data stack and their alpha. They want to know they own the means of production and it's not being transferred to someone else. The current regime does precisely the transfer Karp and companies fear. So he said there are a few things every enterprise must do. Two of Them I'll highlight control. Create your private evals because evals define what good looks like in your organization and then have control of it and then choice. Ensure the orchestration layer is decoupled from any single model model. Ask yourself if any one model you are using is taking away do you still have the ability to operate and optimize your evals using other models? So then he ends. In other words, a company should be able to use a model without giving up the knowledge that makes it unique. That is the reverse information paradox we need to confront. So I mean, as you said up front, Mike, there's just so many related things now each week and like we said with Karp, while he is a controversial CEO and figure if you get, get through that noise, there was a lot of what he was saying that made a lot of sense and that Levy is echoing, that Satya Nadella is echoing. And it is in conflict with these proprietary models that Dario Amade is, you know, obviously a huge proponent of OpenAI is a huge proponent of. But what they're all basically saying is open source is going to matter more and you're going to need more control over the knowledge that you put into these things. And the one thing I will say is to buy or beware. Anytime you're using an intermediary like that, you're allowing your prompts to go through a third party. If that third party is helping you achieve efficiency, figure out what models to use. Like perplexity comes to mind as an example. Here they are 100% taking the data that you're putting in the prompts. You're using the things you're building as training for their own businesses. That is the value to them. It's why maybe they're even going to give these things away. They want your prompts and if you're, you know, if you, if they're an intermediary, it is to extract information from you, to build something else on top of you and your data
A
that's such. And this paradox, I don't know, do we have any sense or advice on how to think about resolving this? Because it's like to get value out of the models, you have to in some fashion provide it with your unique context. I would argue it's not going to happen in a vacuum. You're not going to just magically get value with it out at it, knowing specifics about your business and data. But then to their point, you're giving up your alpha. Is there any. I mean, we talked a bit about some solutions, maybe when we talked about carp, has your thinking evolved at all in that?
B
No. I mean, I haven't had the brain space to really spend a lot of time on it, but there's a part of me that thinks, thinks it's just the cost of doing business. Like, I mean, you could say the same thing every time you use Gmail or, you know, Google Doc or anytime you use Microsoft like that. This has always been the game. The tech companies have always had the insights into every single thing you do, what you click on, what you look at, what you type, what you retype. Like, it's, it's the exchange for the power and the intelligence and the tools. So, so I think it's the hot thing right now and I do think it'll affect some of the ways that, you know, its departments and CIOs in particular think about their use of these tools, especially for more proprietary and sensitive use cases. But I think for like marketers and sales people and CS people, it's like, what the hell are they going to learn from us that they're not going to. It's like whatever we do is public knowledge anyway. In the end, like marketers are putting information. So I, I don't know that it's as critical for the standard like knowledge, work functions. But like if your lawyer is working on patents and things like that, 100%, you should be like on an open source model that you will control internally. That doesn't go anywhere. So I think it, it probably relates more to how you're using these models and what you're putting into them. And again, it's not like it's taking your specific confidential information and training a new model on it, but they're able to learn how you use them, what they use them for, what the prompts you use them, and they're able to, to train the models differently based on how you interact with these things. And so it's like, I don't know, I think it's just for most people it's going to come down to like, it's just the cost of doing business and I don't really care that much.
A
Yeah.
B
Whether that's the right attitude or not. It's just how we've always functioned with business software for sure.
A
All right. Our third big topic this week comes from something we talked about a little earlier in 2025. So in April 2025, the AI Futures project, which is this research group led by former OpenAI researcher Daniel Kok Tajlo, published a report called AI 2027. This is basically a detailed month by month scenario forecasting that AI could automate AI research itself within a couple years, which would trigger an intelligence explosion that they predicted would end in either AI takeover or an extreme concentration of power. This became a widely debated document in AI. It drew recommendations from figures like AI pioneer Yahshua Bengio. We covered this on the show a couple times last year, but now this past week they are publishing a follow up called AI 2040 Plan A. This is a 90 page scenario laying out the team's positive vision for how humanity gets to super intelligence without catastrophe. So the authors are really explicit that this Plan A doc is a recommendation, not a prediction. And kind of like the previous report, they game out like a fictional scenario to show how this could work. So they basically say, quote, we recommend an international deal to avoid a dangerous race to superintelligence. The deal involves total research transparency for AI R D, which allows the nations of the world to understand what's happening and enforce guardrails. The result is multiple companies across multiple countries scaling slowly and safely together towards super intelligence instead of racing each other in secrecy. So basically they're saying under Plan A, if all goes really, really well, humanity could delay superintelligence until 2040. They would make all AI research public. This would allow dozens of companies around the world to catch up to the frontier while safely managing AI development. So we can get into some of the specific details around their predicted timeline here. But this all kind of hinges, Paul, on them predicting like, hey, the US and China basically need to coordinate on coming to some type of deal, whether formal or informal, to start that would actually have them jointly be taking some actions to kind of regulate how AI research works. But they have some real sci fi scenarios in here, like 10 years from now. They're basically saying in this fictional scenario, like most of all labor would be automated by AI, even if they regulate on the way to super intelligence. So there's a lot of very speculative stuff in here and I'm, I'm kind of curious how seriously you're taking some of these. Like I found some of these predictions. I mean, it could happen, but they're pretty wild.
B
Yeah. So Daniel, the author tweeted, plan A is our current best guess, but hopefully a better plan will exist before it's too late. We hope that the best ideas from Plan A will be adopted and the worst discarded. So as I was looking at this Mike, I'm thinking like three to five years out is almost impossible to comprehend. Like the complexity of these models, what happens if we get to AGI and beyond into superintelligence. So this sort of paper can only almost seem like, yeah, like we're not even going to talk about this on the podcast. But instead we chose to make it a main topic. Why did we do that? Well, the starting points actually show why it's so important to work on and to talk about these things. Because what they present and the questions they ask are actually very logical and likely. So they start the report it says in it begins 2027, the writing on the wall. So this is again, I'm gonna. They're presenting hypotheticals. So they're looking into next year. But when I read the start of this, like, this is not far fetched. So the foundation of why they're doing this research rests on this. So America has two workforces now. The first is people, 165 million of them. The second is AI agents. Millions of copies spun up and shut down every hour, working around the clock at superhuman speeds. That is a path we are absolutely on. Most of their work is slop. Definitely a path we were on. But enough of it is good that people are paying tens of billions of dollars a month for AIs that can, in theory at least, do anything on a computer that an employee can do. There is one job the AI companies want to automate more than any other. Their own. They haven't succeeded yet. No recursive self improvement. That's a topic we've talked a lot about in 2026. So no recursive self improvement so far. But they seem to be getting closer and they're pulling up the ladder behind them. The strongest coding AIs refuse to help competitors with AI R&D. That's based on something that's actually happening already. Even as the most bullish employees admit that things are taking a bit longer than planned, the skeptics notice that their usual dismissals are starting to ring hollow. Why exactly will AI never be able to do my job? What's the barrier again? Congress is starting to pay more attention. They've long been hearing about AI data centers using too much water, chatbots encouraging suicide, Mythos hacking NSA systems, and of course, tech industry lobbyists warning that any whiff of regulation will make AI immediately lose the race with China and spend the rest of history as a tributary state, the ccpc, CCP tributary state. So again, none of this is far fetched at all. This is like next year. Now they step back and ask, where are we going with this? What does the world look like 5, 10, or 15 years from now, will there still be jobs? What if there are aren't? One question weighs especially heavily on their minds. Who will control all these AIs? Congress settles on an important part of the answer. Probably not us. They hold a series of tense hearings on AI. They read the 2016 OpenAI emails discussing how OpenAI was founded in order to prevent Demis Asabas from becoming a dictator. But who is preventing Sam or Elon from becoming dictator? Congress is unsatisfied with existing responses. The result of this wakeup is the AI Transparency act of 2027, an omnibus bill that does many things, some good, some bad, but doesn't fundamentally change the situation. So that's what they're looking at saying next year. Now again, is an AI Transparency act going to be created? Who knows? But the rest of that is all very probable. Like it's a. It's a logical thing. So then that leads to 2028. AI is now on the ballots. That's the next presidential elections in the United States. The 2028 election cycle is heated as usual. AI is the biggest topic, which I 100% agree it probably would will be. The data centers now under construction cost twice as much as the entire US military budget. Most white collar professions are seeing disruption. Like software engineering saw in 26. Such jobs now heavily involve managing AI agents. That is absolutely going to happen. AI companies have industrialized the training process. Executives say let's move into X profession this year. That's what VC firms are funding is like. Let's go take on the next profession. And then the company interviews professionals, buys data, creates training environments, etc etc until their AIs get traction. Then the AIs rapidly improve as they are used more widely in the field and accumulate more real world data which then leads to more automation of different industries. Other countries are starting to get scared and angry. It seems like a handful of US and Chinese companies are on track to automate all white collar jobs. Power is concentrating in the US and in particular the President, plus a handful of tech CEOs that's already happening. AI experts warn that the intelligence explosion is near by speeding up AI research, the AIs will become even more competent, speeding up research even faster, making them even more competent, and so on. Both presidential candidates keep getting asked what they'll do about AI and try out increasingly dramatic ideas on the campaign trail. The discourse bounces back and forth across all the options displayed below, which I'll explain in a second. And eventually the President and his protege, which I'm assuming is an AI agent later on Converge on one plan, the opposition candidate converses on another. Then it's election day. So what this is doing is it's setting up. Here's what's probably going to be happening over the next year and a half in the US leading up to this election. And then we are going to have to choose. The candidates will take opposite positions. That's a given in politics. We then as, as the voters will have to choose which candidate we think is best to lead us into this era where super intelligence will likely be possible. So then at a very high level and we won't get into like the, the granular details of the report, they say, okay, race through super intelligence explosion by having AI self improve and putting them in charge of more things, data centers, factories, weapons faster than China can. That is like the option presented to the US and to these candidates. So plan A is a verified slowdown. That's the one you talked about, Mike. President announces that the US will pursue international cooperation to avoid an imminent intelligence explosion. Sounds great. Does not seem like a viable option in my opinion. Union plan B is you fight China. The President announces the creation of a US led coalition to govern AI deployment or development. Plan C is burn the lead. The President says he will be implementing strong regulation to ensure safety and security. Plan D, race to superintelligence. The President says he will be implementing light touch AI regulation to prioritize AI innovation. And plan S shut it all down. The President seeks a global moratorium on AI development. Again, anything beyond 2028 is like completely guesswork. And they're I'm sure relying on really smart advanced AI models as well as their own domain expertise to like develop what this could look like over the next 14 years. But I think just accepting that some version of what they're presenting in 27 and 28 are likely scenarios, it presents the reason why we should be having these discussions now and at least, least contemplating what they're presenting as like wow, we have no idea what it looks like beyond 2028.
A
Yeah, I feel like I had a mild panic attack reading the like setup of this scenario just because like where they extrapolated out to obviously just sounds totally sci fi. But you see the seeds of this being very, very realistic based on what we cover every week. There's huge opportunities and there's a lot of really positive things outlined in their recommendations so to speak. But yeah, it's, it's a sobering read I would say.
B
Yeah, I think you got to be in the right mindset to want to read anything beyond what we just like covered for you. But I do think that it's really, really important, especially Americans like, yeah, it's going to play a role in the midterms in 26. It will, will dominate the 2028 presidential election. Like there's no way. It can't because the implications are so vast across the economy. Energy, where we're getting energy from, where we're building data centers, what foreign entities we're allowing to invest in U.S. companies, whether or not we allow our models to go out, like export the models and chips and like it is going to be a part of literally every conversation that's happening is going to like come back to the implications of AI and the decisions are going to be made. So, like it is going to be very, very important that people understand these, these issues going into 2028.
A
Yeah. All right, before we jump into rapid fire this week, another announcement that this week's episode is also brought to you by the AI for Business Boot Camp by Smarter X. We are coming to Columbus, Ohio this week. When you are listening to this Thursday, July 16th is when we'll be in Columbus for the boot camp. There is still time to join us if you're a leader who's ready to accelerate AI adoption and value creation. This is a single day about 8:30 to about 5:30 at the Hilton Columbus at Easton. We're starting the day off with a State of AI for Business keynote given by Paul. Then we're transitioning into two highly interactive workshops. One led by myself, an AI productivity workshop in the morning and then Paul is leading an AI innovation workshop in the afternoon. So this event is built for AI forward managers, directors, executives across every department who are ready to move past AI theory. We're going to actually, actually work on architecting real AI powered workflows. Get strategic frameworks to accelerate transformation and leave with immediately actionable plans for yourself and or your team. So AI Academy Mastery members get discounted pricing. We also have discounts available for teams of two or more and groups of 10 or more can even get custom pricing. If you are listening to this podcast, you can also use the code POD100 to take a hundred dollars off your ticket. Again, this is happening this Thursday, July 16th. So to grab your spot, just go to SmartRx AI, click on Events. You'll see the AI for Business boot Camp as an option right there. All right, some rapid fire this week, Paul. We had some big news later in the week breaking over the weekend. This past week, Apple sued open AI for trade secret theft Accusing the company and its chief hardware officer of running a coordinated campaign to steal information about upcoming Apple products. So the suit, which was filed Friday in the Northern District of California, says OpenAI encouraged Apple employees to share info, components and other materials tied to unreleased products as OpenAI seeks to potentially build its own AI devices. According to the suit, more than 400 former Apple workers are actually now at OpenAI. So the suit actually names OpenAI chief hardware officer Tang Tan, a former Apple VP of product design who led iPhone, Apple Watch and AirPod developments, along with former iPhone hardware engineer Chang Liu. Apple says Tan solicited details about unreleased products and job interviews. Lou downloaded dozens of confidential hardware files and OpenAI was actually actively coaching departing employees on basically avoiding the kind of quote dreaded walkout when you leave that ends your access to confidential information. So at every level, from members of its technical staff to its chief hardware officer, this suit is saying, and in coordination with business partners, OpenAI has been stealing Apple's trade secrets and confidential information. Apple said. So Apple is seeking a jury trial. They want OpenAI to stop destroy any proprietary materials and redesign upcoming products to exclude Apple's technology. OpenAI's denied these allegations. And Paul, I mean this is a pretty big bombshell. I mean we're going to be watching this one closely. I did not see this one coming.
B
Yeah and just important context. So you know OpenAI acquired Jony I've's company right for like $6 billion or something. So Jony I've of Apple fame. We've discussed like what are the products they're building. We've tried to kind of guess but we know it's hardware related and so we've known for a while that they were working on devices whether it's trying to compete with or replace the iPhone or if it's an ambient listening device device that sits on your desktop or if it's a pen or a pendant. Like we don't, we don't know yet. But they are very aggressively moving into this and I would assume hardware is going to be a key part of OpenAI's IPO like the potential value and market opportunity behind their hardware. The. I'm not a lawyer but holy like the, the, the lawsuit like what they're claiming it is. It does not sound good for OpenAI. So I, I'll just A couple excerpts from Alex Heath who was I follow on on X and he was sort of reviewing the complaint. So he said yeah. Recent, recently significant evidence has emerged suggesting this is quoting from the, the filing individuals employed by OpenAI wrongfully took Apple's secret and confidential information regarding our unreleased technologies, processes and products. We always defend our team's hard work and innovations and we are taking all appropriate steps to do so. Regarding an ex Apple employee named a lawsuit over several weeks while developing hardware for OpenAI Mr. Lou Serendipitous or a certificate? Wait, how is that word surreptitiously? There we go. Yes. Accessed and downloaded dozens of Apple's confidential hardware related files, including voluminous detailed information about unreleased products, engineering presentations, technical specifications and proprietary project data. Other former Apple employees who had gone to work for OpenAI emailed themselves Apple's confidential information to personal account accounts on their way out the door. Regarding Tan, a veteran Apple product leader who is now OpenAI's head of hardware, Apple's investigation has revealed Mr. Tan has been methodically using Apple's confidential information to benefit OpenAI. He has used an Apple internal project code name to ask quote what's the plan for X for an announced Apple product? He has directed job candidates still working for Apple to bring actual parts from Apple to their interviews for show and tell sessions in which he and his team at OpenAI can alert elicits still more Apple confidential information. OpenAI has been instructing Apple employees to bring CAD design artifacts and prototypes to the interviews. In February, an investigation was in the early stages. Apple wrote OpenAI to raise its concerns that Apple's confidential information could be making its way into OpenAI's business improperly. Apple asked OpenAI to discuss what precautions they were taking to avoid this problem, to investigate and to remediate any issues. OpenAI did not respond. Bond OpenAI has been stealing Apple's trade tickets and confidence information. As a result, OpenAI's nascent hardware business now rests on the shakiest of foundations, rotten to its core by its illegal reliance on misappropriated trade secrets. February 9, weeks after he had left Apple, knowing he had no right to do so, Mr. Liu tried to access Apple's network storage. He discovered that surprisingly he could still access the Apple network repository after leaving Apple Apple the result of then unknown authentication vulnerability. Rather than bringing this Apple's attention, Mr. Liu celebrated his fine with Ms. Pang and said about exploiting it quote lol I found I can access the network storage so funny that's not going to play well in court. So it goes on like it is. It's insane. And again not an attorney but like I just this is a bad bad look for OpenAI. Now OpenAI's statement, which I thought was actually a joke. When I saw this I was like, like, but this is from their director of strategic communications. He tweeted our statement in response to this suit, quote, we have no interest in other companies trade secrets. We remain focused on building innovative technology that empowers people everywhere. I didn't realize that was the formal response from Open AI. I was like, oh my God, like their lawyers are going to tell them to get this down in like seconds. But apparently that was their actual tweet. And then Sam said I'm not afraid of Apple, but I have tremendous respect, respect for them. Okay Sam, I again not sure that's going to go so well. So my big questions are Apple and OpenAI are in a partnership. Like they were sharing chat GPT. You can still use it in Siri. So that's probably not going so well. What this does to Apple's or OpenAI's hardware plans like and then the impact on the IPO. I, I don't think when you're going for a multi trillion dollar IPO you want a massive lawsuit for from Apple of all companies hanging over you. So I, I don't know. Like this is really, really interesting and a whole new thing to watch to
A
our point before about worrying that the model companies are taking your alpha. I would imagine at Apple you are not allowed to use Chat GPT today and if you still are, you shouldn't be. It's kind of crazy.
B
Yeah.
A
Oh boy.
B
Well, and that was what led to the Sam Elon tweet that I started with about like him stealing stuff.
A
Well, I'm sure we'll have some updates here soon enough on that. Crazy. Next up. This past week, Illinois Governor J.B. pritzker signed the Artificial Intelligence Safety Measures Act. This is a bipartisan law state leaders are calling the strongest AI safety and accountability framework in the country. This law targets the most capable models built by the largest companies using basically two thresholds if they have $500 million in annual revenue. They also kind of measure if they have massive computing that they're using for the models. And anyone who's covered by this must publish a transparency framework explaining how they apply industry standards, how they measure model capabilities and catastrophic risks and identify and respond to safety incidents. This law also creates confidential reporting channels and whistleblower protections for any employees who raise safety concerns. So Illinois actually then becomes the first state here require regular independent third party safety audits of covered AI systems. That actually goes a step beyond the 2025 New York and California laws that Illinois legislators used as a Model Illinois Attorney General Kwame Raul will have authority to find companies up to $1 million for a first violation, up to 3 million for each additional violation. This law passed the Illinois House 110 to 0. And it actually, actually did draw support from OpenAI and Anthropic. Anthropic actually said it was proud to be the first AI lab to support the bill. So Paul, I'm curious what you think of this law. Like we're starting to see, it seems states fill the void, which is kind of what the federal government, at least the administration was worried about. If there's no federal overarching law, states are going to pass their own regulatory frameworks, it sounds like.
B
Yeah, I didn't, I didn't see any response from the government on this one. The federal government, government, I would imagine they're not like huge fans of this and they probably aren't fans that openly anthropic are publicly supporting it. I don't know if this is like a skeptic in me, but I just find this hard to believe it's going to have any real impact. You know, if Open and Anthropic are like, yeah, this is great, then it's like, okay, yeah, it's probably, it doesn't have much teeth and I don't know if it's like that. It's only like a 3 million dollar violation. For each additional violation, they're like, yeah, okay, that's like a rounding error. Or that it really just isn't that aggressive. I, I don't know, it just doesn't seem like it got a ton of pushback from anybody, which tells me it probably isn't like a huge deal. Yeah, but it is a sign of states making progress and filling that void, as you said. But I, I don't know that it's like a, you know, this is a transformational thing when it comes to, you know, how the models behave and what they're going to do and the power that they're going to have. So, so I don't know, I might be wrong on that, but it just doesn't seem like it got a lot of run. And that tells me it probably isn't like a massive deal yet.
A
And like we talked about with the California laws they were considering, like, good luck keeping a close eye on like compute levels and where the thresholds are. Like, I don't even know how you begin to do that at the state level.
B
Yeah. And again, like one of the, one of the loopholes that I just seen, it seems like we realized is going to be an ongoing issue is all of these regulations are related to publicly released models. So the loophole is they don't have to release the most powerful models. The labs can have more powerful models. They can give exclusive access to select companies to government. And so this doesn't cover the fact that the model is just going to keep getting smarter and like maybe we just start reducing who has access to them because it's too hard to release them publicly. So they just do these limited releases and then the general populace doesn't ever get the most powerful models. I don't know. But like that seems like a logical path that all the regulation could lead to. Yeah.
A
All right, next up, something else related to AI safety. This past week the Future of Life Institute published their summer 2026 edition of their twice yearly AI Safety Index, which convenes an independent panel of seven AI experts. Grade 9 leading AI companies across 37 safety indicators. Unfortunately, the top grade was a C. Anthropic again held the top spot. With that C open, AI slipped from C to a C grade. Google DeepMind ranked third after that. The panel noted all three have weakened or dropped earlier pledges to halt development on their own if certain red lines came into view to they have also softened their resistance to military uses of their technology. So Meta actually improves slightly. Climbed from a D to a D plus. It'll sound great. Good for you. Elon Musk's XAI obviously just rebranded to Space xai. It fell to an F. It joined China's Deep Seek and France's Mistral at the bottom. The institute's co founder and president Max Tegmark, we've about talked talked about before, said the failing grades spanning three continents show that this is a global problem. He Stuart Russell, who is also a panelist, said companies have backed away from earlier commitments to release new systems only with safety measures appropriate for their capability levels. Now they're planning to release them even if it's demonstrably unsafe to do so. Tag Mark told Time magazine that a real race to the top on AI safety will take regulation, saying he is cautiously optimistic. Optimistic and pointing to the new the EU's AI act, new Chinese rules that are taking effect this month, and a more risk conscious US administration. So Paul like no question, Tegmark comes at this from a very specific view on AI safety. But him, Russell, the other panelists are pretty big people in AI saying that the companies are not doing a good job. It sounds like.
B
Yeah, I don't think that's surprising to anybody I mean, certainly, you know, some of the labs do a better job job of releasing information, being a little bit more transparent, what they're doing, being more vocal about the risks related to what, you know, everybody's building. But some of these labs are, you know, very intentionally not sharing this kind of information. Yeah. Like, so that it's. I do, I do think it's funny though. Like, the D to the D plus is really good. And but again, if, you know, I know some of our listeners are really concerned about the AI safety side of this and where does this all go? And so this is just for you to know that there is a report out there that like anything in politics or in AI, there's, there is bias in the Future of Life Institute. Like, what do they believe and what are they pushing? Like, you always have to put the context of, okay, who is publishing this, what are they traditionally focused on, what is their mission? All that being known, though, like, it's, it's important to, to know these things exist and be able to go down this path if you want to go read this stuff. And there's nothing in it that surprises me. Like, we know that they're not really doing a great job in this area, but it does sort of quantify it in some ways.
A
Yeah, it really does. All right, so this past week we also saw a cheating scandal at Brown University that became a viral case study in what AI is doing to higher education. So this is about an economics professor, Roberto Serrano. He basically allowed his students to take take home exams in his very difficult welfare economics course. This was the first time had done it this spring. Unfortunately, it was due to, because in December they had a campus shooting that left a bunch of students anxious about being in classrooms. So he said, okay, we'll do the exam as a take home this year. Interestingly, when he announced that enrollment jumped from a typical class of under 30 students to 86 students, the take home midterm came back with an average score of 96 out of 100. 40 students scored a perfect 100. The historical average in the course range between scores of 65 and 80 on this exam. And he basically noted this. He saw that many answers had this weird convoluted style. So he started running the exam questions through ChatGPT and produced basically similar answers. So what he did is he said, look, we're moving our final exam to in person and if the score distributions look the same, we'll hold the midterm grades where they are and you can take your, your well earned grades if the score distributions look different, we're voiding the midterm. The moment he said this, 18 students dropped the course. Nine more skipped the final 22 of those, 27 had scored a perfect 100 on the midterm. Here's the thing. The students then took the final in person, and the average score dropped from 96 to 48. And so basically, while he says he's not using AI detectors or anything, he's like, look, this is a pretty solid indicator. Everyone's chat, GPT or other AI tools to basically totally cheat on all this stuff. And so it's kind of set off this scandal where the teacher's like, look, you know, he softened the blow a little bit by changing some of the worst scores, it seemed, or given a little more credit. But he's like, we've got a serious problem here. And so, Paul, I don't know if this is surprising per se, but seeing those exact numbers was crazy to me.
B
Yeah, that's the. It's jarring just to, to see, like, it quantified like this. I don't think anybody would be surprised that kid, kids are using these tools to cheat. So I like the rapid fire. I won't go deep on this, but the thing that jumps out to me is you can't outsource thinking. And we have to, as employers, as parents, for ourselves, for co workers, we have to be constantly aware that we have access to intelligence on demand. And if we use it as a crutch for everything, we will forget how to think. And, and so this becomes critically important, obviously, for educators, that they're, they're not allowing, they want to teach them how to use AI responsibly, but you cannot do it where they just outsource the important part of what we do. And you have to think about this when hiring and when evaluating your existing employees. So in hiring, you almost have to assume that the people you're interviewing out of college got through college using AI now, right? That can be really good. But you have to be able in your interview process to test for critical thinking capability to, to make sure that they didn't atrophy that ability in college, that they've lost the ability to think for themselves and to assess things. So it's really, really important. Like, this is. The whole next generation of workers will have never known education at the higher levels where they didn't have the ability to do it, and they probably didn't have someone over their shoulder making sure they weren't using it to replace critical thinking. So that could be a problem I could compound within a company.
A
Yeah, this feels like the real danger at the heart of all the positives about AI usage for students especially. It's like we're seeing it in the corporate world with what we've called in the past, like work slop.
B
Right.
A
Where people are just submitting AI responses to their co workers that are just not nonsense or like, don't have any thought put into them and no, makes more work for everybody. And it's like, I just. If you're a student and you think that you can come into a job and just submit whatever ChatGPT gives you, I don't know if you people think that, but I would just say you have another thing coming if you think that is the standard moving forward.
B
That's for sure.
A
All right, so another big story this past week, the New York Times reported that state actors in China, Russia and to a lesser extent Iran are working to inflame American debate over AI data centers. According to some analysis they're reporting on from the threat intelligence firm Alithea, between January and June, state media in the three countries mentioned data centers roughly 700 times, an average of nearly four times a day, basically in an effort to turn data centers and the controversy around them into what Alethea calls, quote, a domestic fracture point heading into midterm elections where AI is seen as a top issue. So there were some examples that included a Chinese state owned newspaper publishing a satellite image of data center in Gainesville, Virginia, warning that AI threatens Americans. Well, being there is a chat GPT generate comic strip that was misinformation disguised as a Maryland news outlet blaming data centers for soaring electricity bills and a known covert Russian operation circulating a video attacking an American company's data center project in Armenia. So, Paul, we have known for over a decade at least that foreign adversaries of the US have routinely pushed misinformation about tons of different subjects over the years to basically, basically stoke domestic turmoil. It seems to be happening with data centers. I believe this is yet another podcast prediction that came true or was correct is that this was happening. It seems like we have evidence of at least a little of it. It's not something where it's like every single thing you see. There's a real invalid debate around the issue. But it sounds like some people are trying to take advantage of that.
B
Totally. Yeah. I was like, I think when I first mentioned it like a month or so ago, I was trying not to go down this rabbit hole too much. I mentioned like Cambridge Analytica and things like that. Yeah, this is absolutely a strategy that foreign governments use to influence citizens in other countries. The US does it to other countries and people do it to us. And so if you find something that causes friction within the citizenship, then you, you push on that and you use social media and now you can use AI tools to personalize the stuff that you know, pushes these ideas. So it's not to say, as you alluded to Mike, that this isn't an issue, that there aren't real issues, but foreign governments are absolutely going to push on these things to try and create a more, I guess, hate and fear and distrust among American citizens. So it, it's just a, like people need to be aware and maybe you need to like educate your kids. Like if your kids aren't aware how social media works and things that they're seeing on TikTok and YouTube and wherever else they get their information that, you know, sometimes it's not real and it's actually meant to piss them off and get them all worked up about a topic. AI is just one example of this, but like your kids need to know how the world works for better or for worse. I've had this conversation with my 14 and 13 year old, so I don't know what, when's too young to start them and my kids aren't even on social media media. They don't have accounts like they, they use YouTube but they're not on TikTok and Instagram and things like that. And we've already had deep conversations about how this stuff works.
A
Yeah. Wow. Okay, next up we have our AI use case spotlight where every week we give you a quick look under the hood at some real AI use cases we're exploring. So I was just going to share one quick thing, Paul, that came out of my prep for our AI productivity workshop at our boot camp this week. So as part of that I was building out categories of AI capabilities basically to help participants understand what can AI tools do today that you might not be aware of. So I actually used Codex to build a huge source backed like capability spreadsheet across Chat GPT, Gemini, Claude and Microsoft Copilot. So basically the way I did this is like described what I wanted to do ideally. And Codex split this research across four separate agents, one per platform. So one agent researched chat GBT using only official OpenAI sources. That was a specification. Same thing for Cloth, Gemini and Copilot. Each agent was asked to inventory the current end user and business capabilities of each platform as of July. At this point it was July 8th or 9th 2026. So things like the models, the reasoning, chat search, web grounding, deep research, et cetera. Like what are all the things that are today available in these tools? And it created. It took like over an hour. I think this is like probably one of the longer use cases I think I've had for in a while. The final CSV has 177 rows do each one documenting a feature, showing what link and the documentation it came from. So it was like 200 roughly for each of these. We then did another pass to audit it basically looking for blank fields, looking for missing source links, stress testing all the answers. Super interesting. It's deeply overwhelming. It's not like useful as a public facing asset I would say. But. But it was extremely critical in distilling this topic down into something others could understand and making sure that was all backed by real data. So it was awesome.
B
That was really cool. I did a. I shared this with you Mike, so you've seen this. But I actually was trying to think of something to test Fable 5 with since they extended access to the 12th and so I think it was like Friday morning or something or I guess no, Saturday. I'm looking at it now. I ran this on Saturday, Saturday morning. So I'm not a huge one on like tracking competitors. I don't, I don't. I generally just like focus on listening to our audience, looking at the trend data and like building what needs to be built what we think you know is best. But every once in a while it's good to just like see what's out there and see what's going on. So I actually ran a competitive analysis and I gave this to chat. GBT 5.6 Soul High Work Edition it looks like. And then I also ran it on Fable 5. So I'll give you the, the exact promise that run a competitive analysis on and then the name of the competitor. Consider strengths, weaknesses, threats and opportunities in comparison to our business. Propose business strategies that we can use to exploit their weaknesses and our strengths to differentiate in the market and be the clear choice for enterprises. So that. That was it. That's the entire prompt. And then I gave it to both of them and it, it crushed it like they were. You saw them, Mike. Like I and I. What I didn't this case again, like the whole idea of AI slop. I shared them internally with a couple people and I said listen, this is. I don't have time to edit this. I just wanted to run this as a, a test in these two different models. I've taken both the outputs, I've put them into this single Google Doc. This is the raw unedited version of this. I'm not going to get to it until later next week, but just wanted you guys to have access to this as well. And so I get like, I think one, just the simplicity of the prompt, two, using it for high level strategy. Three, if you're going to present something that you haven' verified yourself and given the time to think about, don't present it to your co workers as though, here's this genius thing I did. It's like, no, here's something I took 35 seconds to do across two platforms. Here is how I'm going to move forward, verifying this and using this information. But here's the RAW files in case you guys have a chance to take a look at it or it's relevant to anything else you're working on. So that was a cool thing and I'm really anxious to actually dive into, into that later this week.
A
Yeah, the outputs were so cool. And I would just mention, as you're kind of talking about that it really reminds me of what we talked about at the top of the episode that Ethan Molik was saying that this kind of strategic knowledge work is not coding. So like you could spin up a bunch of agents to like stress test these ideas. It might redo, result in something really useful. But you would need to then go audit like what the logic was there. Right. So it'd be much more you working in tandem with this raw output versus is, hey, let me turn an agent loose and have it debug this thing. Right. So it's kind of very interesting to see how this knowledge work differs from coding.
B
Yeah. And the other thing is you can then take this and say, okay, you know what, I'm really intrigued by what they're doing every Friday morning. Run an updated report, tell me anything new they've done and now you can get at the agentic side and start again. Changes work. You reimagine what it's like to do these things. Things like, Mike, when we were in my, like the agency days, you and I, I mean, we would charge like, I don't know, 10 to $20,000 to do these like deep competitive analyses and then provide strategy on top of it. And it did it. 28 seconds. It's wild.
A
It's incredible.
B
Yeah.
A
All right, we're going to wrap up here with some AI product and funding updates I'm going to run through real quick because we have a ton of them. So first up, OpenAI published an approach to government and national security partnerships laying out principles for its growing defense and public, including commitments not to allow its technology to be used for mass domestic surveillance, direct autonomous systems or make high stakes automated decisions. OpenAI CEO of Applications Fiji Simo announced she is actually stepping out of her role. She's stepping back to a part time advisor role. She's been on medical leave for a chronic illness for three months so she will no longer be in that role. Anthropic has as of right now once again extended access to Fable 5 on all paid plans through Sunday, July 19th. It was originally supposed to end end July 12th, but now anyone who has a paid Plan can use Fable 5 within certain limits rather than only paying use it for usage as you go. So when you're listening to this, you have a few more days to try it out. Anthropic appointed former Federal Reserve Chair Ben Bernanke to its Long Term Benefit Trust, the independent governance body that oversees the company's public benefit mission and can appoint board members. They cited how his ex his expertise on how AI will affect affect workforces and economies. Anthropic also brought Claude Cowork, its agent platform for delegating everyday knowledge work to web and mobile, letting users hand off a task at their desk, monitor progress from their phone and let work keep running in the background. Microsoft has started replacing OpenAI and anthropic models with its own internally built Mai models for some AI features in Excel and Outlook. Meta released Muse Spark 1.1, an agentic encoding model, and began charging for access to its models for the first time through the new Meta model API pricing that CEO Mark Zuckerberg says the pricing is actually roughly a quarter of what Anthropic and Open AI charge. Funnily enough, he announced this in his first post on X since 2023. Meta also launched Muse Image and Muse Video. These are new image and video generation models. They immediately drew privacy backlash after the New York Times reported that public Instagram accounts are enrolled by default in a feature that lets other users generate AI images from their photos. The Financial Times also reported that Meta is testing prototype AI glasses designed for continuous ambient recording that feeds an onboard AI memory users can query. Meta has reportedly considered whether the indicator light would actually stay on during passive data collection. SpaceX AI, the new merged company formerly known as XAI and SpaceX, is rolling out Grok 4.5. This is their new Frontier model, built for work beyond software engineering. They're rolling this out to all customers after a beta test program. And Cursor is making the model available across its coding platform. Amazon is working on a secret project codenamed Moonraker to turn Alexa into an AI agent that can chain together multi step tasks from a single command with internal documents, projecting more than $100 million in GPU costs. In 2026 alone, a startup called Prime Intellect raised a $130 million Series A LED by Radical Ventures with backers including Nvidia and Intel to build what it calls the open Super Intelligence stack. And Artificial Analysis introduced six new capability indices for comparing AI model capabilities capabilities across industry domains. And I think Paul, you had one or two more things to highlight as well.
B
Yeah, this one just dropped this morning, so I'll just throw this in here. So this is from the Stanford Digital Economy Lab. There's a new website. We must act now. AI is the URL. 16 Nobel laureates join leading economists and AI researchers and call to prepare for AI's efforts. Economic transformation. Calling for urgent preparation for the economic impacts of radically more powerful AI. This is led by Eric Brson, Ajay Agarol, Anton Cornick and Tom Cunningham. But it is also signed by Noam Brown, who we talked about Jeff Dean, Jack Clark, co founder of Anthropic Shto Douglas. We've mentioned Yasha Benjo, Eric Schmidt, Dean Ball, Ben Bernaki, who you just mentioned mentioned. So the statement is three, it's three simple points. One, AI may become radically more powerful over the next 10 years. Two, this could drive an unprecedented transformation of our economy, larger than the Industrial Revolution but unfolding over a vastly shorter time frame. It could bring risks, including large scale job displacement, as well as opportunities such as major gains in living standards. And three, Economists, policymakers and technology leaders must act now to understand the economics of transformative AI and to build the incentives, guardrails and institutions needed to steer AI in a direction that complements humans and benefits society.
A
Seems like at least a few economists are waking up to what we have been talking.
B
Yeah, and it's interesting because it's like there's, we've mentioned recently how the lab, sort of an apparently coordinated effort, all of a sudden started backing off of their belief. Like Sam Altman in particular, even just tweeted again on Sunday like, oh yeah, I might have been wrong about jobs. It's like, no you weren't. Like it's coming and like they're all, you know, starting to accept it. And I, I think, I don't know what the other goals behind this are. We'll put a link to the announcement from Stanford and a link to the site. But it's, it has to happen. Like I've always said, at minimum, we need to be prepared. If it doesn't happen, great. But we need to be prepared in the event that it does cause massive displacement and under employment.
A
All right, so that's a wrap for this week. Paul. Just one more quick reminder to visit our AI poll survey for this week. SmarterX AI Pulse. This week we are asking about your feelings around ChatGPT work and also your feelings around using Voice AI. So if you go to that link, SmarterX AI Pulse, we'd love if you shared your thoughts. Paul, thanks again.
B
Yeah, programming Note no Episode July 21st first so, right, so I will be on vacation and we were going to try and squeeze in a recording early, but Mike and I are in Columbus for the AI for Business boot camp this week so it just, it wasn't going to work despite our best efforts. So we're going to one week break for the podcast. We'll call it our summer break. Mike give you a day off?
A
Yeah, there you go.
B
And, and then we will be back on. I guess that would be what, July 28th?
A
Yep, yep.
B
All right, so thanks everyone. Have a great two weeks, I guess. And we will be back with the Latest on episode 226. Thanks for listening to the Artificial Intelligence Show. Visit SmarterX AI to continue on your AI learning journey and join more than 100,000 professionals and business leaders who have subscribed to our weekly newsletters, downloaded AI blueprints, attended virtual and in person events, taken on online AI courses and earn professional certificates from our AI Academy and engaged in the SmartRx Slack community. Until next time, stay curious and explore AI.
Episode #225: GPT-5.6, ChatGPT Work, Enterprise Agents, AI 2040 & Apple Sues OpenAI
Hosts: Paul Roetzer & Mike Kaput
Date: July 14, 2026
This episode delivers a comprehensive breakdown of the latest AI advancements, focusing on OpenAI’s GPT-5.6 launch, the introduction of ChatGPT Work, the rapidly changing landscape of agents in enterprise, new regulatory and legal developments (including Apple suing OpenAI), and strategic perspectives on AI’s societal impact, including a deep dive into the new “AI 2040 Plan A” scenario. The hosts aim to give professionals actionable takeaways, insight into where AI is headed, and the strategic frameworks necessary to navigate the AI-powered future.
[05:46–13:00]
GPT-5.6 Model Family:
Launch Complexity:
GPT Live Voice Models:
“There are a lot of benchmarks that suggest 5.6 SOL is the best model in the world right now, but the most reliable way to tell is that Elon is obsessed with me.” — Sam Altman [11:53]
Elon: "...what do you plan to do for an encore that's tough to beat?”
Paul: "So it’s always good to have the Sam vs Elon soap opera back." [11:53–12:16]
[13:00–25:51]
Overview:
User Confusion:
Desktop App Differentiation:
[22:27–27:11]
[31:30–46:39]
Aaron Levy (Box CEO) on “Operating Model Problems”:
BCG’s 2026 Survey:
Satya Nadella’s “Reverse Information Paradox”:
Bottom Line:
[46:39–57:45]
[57:45–66:02]
[66:10–71:57]
[86:19–92:25]
For Businesses:
For Individuals:
For Society/CIOs/Government:
For further reading and supporting links (including to detailed reports, product announcements, and the AI 2040 scenario), see show notes at SmarterX AI.