IBM’s cloud collab, Meta’s Muse Glimmer & OpenAI’s upcoming Astra model

IBM TechnologyPublished Aug 14, 202636:33Added Sep 7, 2026

Visit Mixture of Experts podcast page to get more AI content → https://ibm.biz/~YOqgkKvh7

Watch on YouTube →
Contributed by Heather

Transcript

Transcript format
Chapters4

Intro

00:01The best way of solving this is open is always the best way forward. And I think at some point and I'm going to call it and I've called it before, there will be a point where open will overtake the closed models. That will happen at some point, right.

00:15It's just a matter of time. And when that happens, we're going to we're going to have some fun. All that and more on today's MoE. I'm Tim Hwang and welcome to Mixture of Experts. Each week brings together some of the best brains and artificial intelligence to lead you through the week's news.

00:34On this week's episode, we've got Chris Hay distinguished engineer Rynne Whitnah, technical lead AI ecosystem, and Volkmar Uhlig, CTO and VP of Data Platforms. We're going to cover three big stories today. I really want to talk about Muse Glimmer, and that is new model.

00:46We'll talk a little bit about Astra OpenAI's new model. But first I want to start by talking about a big new partnership that's just being announced between the and IBM and.

IBM AI Partnership

01:05IBM's going to be teaming up with Together AI and Nvidia to launch and make accessible the B 300 generation of chips to the public. And so I know you've been pretty close to this. This is a pretty big story because, as I take it, you know, IBM will basically be in the game on providing sort of this next generation cutting edge for compute.

01:22Do you want to take a little bit behind the scenes, like what does it take to have this sort of thing happen? And I guess why Together as a partner here. So Together is one of the bigger AI vendors providing training and inferencing capacity.

01:35And those systems are very differently designed from traditional cloud computing environments, primarily because they need networks which are super high speed. You have a reliability issue. So there's a lot of build up around redundancy, so that if individual node fails, then the overall system stays alive.

01:55Do you need to upgrade ability etc.. But then also there's a big build out on power infrastructure or cooling infrastructure and overall network capacity storage. So we are building out these systems in IBM Cloud. We have been having them for internal use cases or small scale use cases.

02:17Internal. We built pretty large training clusters and now we are entering the market and we are having a partnership with Together AI. Which brings us to the to the end customer and to to providers in that race. So I think what we're I see this moving to is we're in a transition right now from, you know, AI is kind of this esoteric thing to industrial scale computing.

02:39And industrial scale computing requires these industrial scale deployments. And if you're seeing this across all the hyperscalers, we are going from a phase of like, oh, this is this esoteric piece of hardware and kind of cool tool to know.

02:55Now we need to go and get into the cost optimization. This is human infrastructure, right. And so I think there will be a few more players in that game, but it's a very capital intensive like industry to be in. And it's also technologically hard to do this at scale, putting, you know, GPUs under your desk is a different problem than putting 10,000 GPUs into a data center.

03:19And so this is where the work with Together AI. And we're building on years or a decade almost of AI deployments for internal purposes. And, you know, now we have this accessible at very large scale together. Yeah. Together AI. And I think industrial scale I think is the thing I want to pick up a little bit on.

03:39You know, Chris, you know, over the last few years we've seen the emergence of these so-called neo clouds, right, that basically come to the market saying we know how to do this. Specialized AI, you know, compute development. And this is why you should go with us, and you shouldn't work with those old school clouds because they just don't know how to design these kind of special types of clusters.

03:59And I guess, Chris, I guess maybe the question for you is how long do you think that's going to be the state of play? Because I think to Volkmar's point, it does feel like, okay, now, now the really big boys are moving in, the clouds are going to really be kind of doing this at mega scale and are building a lot of competence in it very quickly.

04:15You know, are neo clouds going to be able to survive in this environment. Is there going to be a niche for kind of specialized AI cloud providers, or is the future of AI cloud just kind of the cloud? If I look a little future forward and I look at what everybody is doing, I think there is space and is probably not for the reasons you think, right?

04:33Because if we think of the the large industrial cloud providers, they're investing for a very long time, right? You know, you're you're going to have these GPUs going model architectures are going to change. And therefore you're going to have to be able to handle the new model architectures as they come out and be able to run at at speed and scale and be able to have that running for the large enterprises and be consistent.

04:56The neo clouds, I think, are under less constraints. In that sense. They can be a little bit more innovative and they can go a little bit faster and they can have more agile architectures. So I think that I think there's still a place there, and the two are going to sort of coincide.

05:14What I think is probably going to get super interesting is we're already seeing that most of the AI providers, their demos these days when they release a model is, look, I designed an AI chip, do you know what I mean? Now, I know I don't. I personally, I don't think that's a coincidence at all.

05:34I think that's something that everybody's thinking about is like, how can I use AI to design more efficient chips and then be able to run my own chips, especially for things like inference? So I, I probably think that actually there's a lot of diversity that's going to come in the future.

05:47Yeah. It's almost that's really interesting question between basically like do the do the models have the leverage here or does like the kind of compute and infrastructure sort of have the leverage? You know, I feel like there is this kind of like dream of vertical integration that certainly I think some of the frontier AI labs have, which is okay, well, you know, we're going to just build our own infrastructure, we're going to design our own chips, and it's going to be specialized for our models, and life is going to be great.

06:09I also remember, I mean, back in the day, I was shocked to learn that Netflix ran on AWS for a very long time, which is like a lot, a lot of bandwidth, all provided through a third party, I guess, on that part of the market. Not even talking about neo clouds, like if you were an Anthropic or an Open AI or one of these frontier labs, do you think the idea of kind of vertical integration here is probably something that you'll get to at some point?

06:36Or, you know, again, just in practical terms, you will have to offload this to some other cloud provider. I think it's actually both. And the reason I think it's both is because I think a lot of the chip plays right now are a diversification play, where you need to be able to show that you're growing, especially as you're nearing IPO, and especially as you look into going into the future and with how much difficulty we've had getting RAM and things like that over the last couple of years, you have to have your own supply chain if you're going to preserve yourself from some of those risks.

07:11So I think that's been a lot of the play. But of course, also efficiency of data centers is a major cost advantage. And if you can have something that can run your models at some improvement to that, it's going to result in more profit. Even if you design, you know, a chip and you're suddenly figuring out, oh, this was not an optimal design.

07:36Those hardware investments, if you don't outsource it to a third party, is on your books. And so I think what you're doing is at some point it's a chicken and egg problem. Right? So at the beginning, like, okay, you're designing your model for a specific piece of hardware, then you're adjusting the hardware to your model.

07:53And then at some point you have these assets on the books and you will have to run them out for five years. Otherwise you kind of recoup your investment, and then you will actually design your models for your fleet. And this can happen when you control the supply chain, not if Amazon controls the supply chain for you.

08:11Right. And then the other thing is Amazon. You know, they're running currently at 43% profit margin if you look at AWS. And so I don't think that if you look at the token economy that the combination of infrastructure running and model running allows in the long term, not in the short term and long term, allow that Amazon takes a 43% profit margin right now.

08:36The reason why in many cases they do that is because elasticity is allowing them to, you know, the users of this infrastructure allows them to scale up and down and shift the demand and the elastic demand curve onto the cloud providers. But the cloud providers therefore make 43% with this elasticity.

08:57And so I think we are getting into a point where the consolidated factory of AI is while I use my machines to run inferencing, I may use them for training. And training doesn't necessarily mean like, hey, I'm fine tuning your model with reinforcement learning.

09:15A lot of the training is actually inferencing, and so I can use my inferencing capacity for training phases, etc. for back testing. And so I think having the integrated factory and owning these assets and being able to shift them actually makes economic sense because you have a baseline load.

09:34And so the elasticity Amazon gives you if you are at that scale, it's just not beneficial anymore. So at small scale it is but at large scale you could just run your own much more efficiently. And since it's such a specific use case, you're running 99% of your workload good enough, right?

09:49So and then you use Amazon for everything else or Azure or GCP or IBM. Yeah. Yeah. I think there's a really interesting question. And maybe Volkmar, if you want to have a final thought on this is, you know, there's a really interesting question here.

10:02Just basically like how spiky do you think the median enterprise user's consumption of tokens is over time? Because it totally influences the kind of shape of the cluster, I guess, in your experience. I mean, what are you seeing in terms of like industry trends and obviously varies, but yeah.

10:18Yeah, what we saw is that, you know, I was wondering what's next. And so we saw these load shifts. So you kind of have a base load which is global. And then in an individual zone you have about three x difference between day and night cycle.

10:31And so it's quite substantial. So you know in the fluctuation. Now on the flip side, if you look at the worldwide load it's actually much more the fluctuation is about 20, 30%. And so if you can if you can load shift between the different data centers, you can actually even it out.

10:53So now there is a problem which you know, when you're in highly regulated industries like IBM's customers are. Sometimes load shifting is really, really tricky simply because you have data sovereignty requirements. And so what you see now is all these providers, they are doing the load shifting to batch processing.

11:12And so what you're doing is you're taking the interactive workloads during the daytimes. And if you look at all the guarantees you get for batch processing, it's usually 24 hours. So you know, you submit your batch job. And so what you're doing is you're using the batch jobs to backfill the load fluctuations, and you're just time offset them.

11:30And that's how people, you know, kind of even out their GPU utilization. Well great. I'm going to move us on to our next topic.

Meta Muse Glimmer

11:46It's kind of been this last week, sort of the the return of the Zuck. You know, Meta's been quiet for some time. They had a huge news cycle about all the talent they're recruiting. And then kind of just there was sort of rumors swirling for a bit.

11:58And we are starting to see kind of a bunch of announcements come out of Meta. It feels like they're really kind of cranking, you know, the press engine on their releases. And this past week there's two things that happened. First one was the launch of a new model.

12:12So Meta has launched something they call Muse Glimmer, which is what they bill as kind of an open model that runs on device. And then separately, there's been kind of a vision piece, you know, a little bit like Sam Altman's, like the Age of Superintelligence.

12:27Mark Zuckerberg has also released his own sort of essay called The Future Is for Everybody, about kind of the desirability of distributing superintelligence to everybody. And I guess, Chris, I know when I mentioned that we were going to cover this story, you had a big grin and you were like, oh, I've got some things to say about this.

12:42So how about I just let you kind of sound off? I'm curious about, you know, how you received sort of the post and I guess what you think about the model, what's your what's your taste test on the model? So the first thing is I'd like to welcome Mark Zuckerberg who's clearly a listener to this podcast, because I don't know if you remember this two weeks ago or three weeks ago, it was Tim.

13:04I did say we talked about Spark, and I was like, I don't care. It's like, open up your models, you know, nobody cares if. But the good news is he obviously listens. And he was like, yeah, I need to open the model up. And he did. And suddenly I think I did.

13:16Sharp Guy. The new model actually on the Glimmer model let's talk about that for a second is great. It is actually great. There is if you if you haven't tried it go download. I honestly think it's one of the best small models that you can run.

13:32It's a 30 billion parameter model. It's a dense model, which is nice to see something. It's not a mixture of experts for a while. They're they're clearly competing with Google. So it's really nice to see architecturally they're taking a lot of the insights from like the Gemma models.

13:53And they're taking insights from the Phi-3 models. For example, it runs beautifully fast. I was running on my Mac M3 and it was running. It was running at around 30, 40 tokens per second, which was pretty reasonable. And again, I think on my M2 it's a lot slower, so ten tokens per second.

14:16But actually it is a great model and they've brought a lot of fantastic insights. So they're doing speculative decoding with their flash drafter, which is basically they essentially predict in blocks. You know what the next tokens of the models are in parallel.

14:31And then they check it back with the model going correct. So there's a little mini model doing that draft. And if it's correct then they accept it and they move on. And which is why they're getting speed. They've clearly optimized the model down for the sort of 24 gig, you know, like essentially your Spark boxes etc..

14:51So they've really worked on worked on that. They provided the quantization levels. They've really thought about it from a model architecture perspective, and then they've spent a lot of time on the context as well actually. So they're sharing Google's techniques.

15:07So they're I'm not going to go too much into the technical details of this, but effectively, when you're looking at the residual stream, I'm not going to go into the details of that. But basically they're spending most of their time looking at small context, the last 2000.

15:24But every so often maybe I think it's every fifth block or so. They are then looking at the context as a whole. So they're getting the benefit of being kind of looking hyper focused on your last set of tokens, but every so often just keeping an eye on it.

15:38And they're techniques that are being used in the Gemini models. Their techniques are being used in the Gemma models, for example, which they're clearly competing with. But it turns out what you're getting is the super fast, super small model that runs on your laptop that's optimized and designed for your laptop.

15:51And it's clearly designed for agentic, right. They're focused on research. They've focused on tool calling. It is honestly a great model. And again, if you compare what I said a few weeks ago where I didn't care, now I care. Well done, Mark Zuckerberg.

16:09Go write as many essays as you want as long as you keep releasing models. Instant Meta fan I think I want to take it like this discussion maybe on just a tour of recent history, right? Because if you recall, Meta used to be really leading the way on open.

16:24Remember every Llama release being like, oh man, the new Llama released. This is huge. It's like, you know, and then Meta kind of faded a little bit. And I think the narrative was, okay, well, these kind of Chinese labs are really leading in sort of the open space.

16:36I guess the question is whether or not we're like, kind of overly panicked about sort of like, you know, the Chinese lead in open, right? Because I think we've got Gemma now. We've got Glimmer. These are really good open models. And I guess the kind of question is like, you know, I guess maybe two things.

16:51One of them is this kind of like, does this put sort of Meta back in the running, but also maybe the question of like, actually, maybe like US labs are really good at open as well is I think what I'm reading from an announcement like this. Yeah.

17:00I mean I'm going to throw one more in the mix. This has been IBM's focus around the Granite models, right? So generally a big fan of this approach, right. Where, you know, the more we can get in the open, the faster we're going to innovate, the faster we're going to see things changing.

17:19And that is, of course, as always, a double edged sword, right where there is now the ability to run these agent workloads on your laptop. And that's something we can't undo, right. That's now just baked into the future. Right. So as we look at all of these models that are breaking containment and things like that, as they start having security vulnerabilities in their sandboxes and such, I think this, more than anything, solidifies AI development as a requirement where you now need to be able to use top of the line agents as part of your coding,

17:56because everybody else has them as part of their threat model. So all of those things come together to say that, you know, yes, open weight models are the solution to that. That's how we get research. That's how we continue developing. That's how we push the field forward.

18:09Running them on your laptop also critical, right. Like that's how you keep costs down. That's how we keep from everything being in a data center. And that's how we democratize the tech. But everything has a shadow side right. Do you want to talk a little bit maybe about the kind of compute and infrastructure aspects of this?

18:27Because obviously, like the big headline that's trying to link to these stories is that Glimmer is sold as a model that runs on your device. And Zuck, of course, is sort of selling this vision that superintelligence is going to be sort of everywhere and for everybody.

18:41And it's actually kind of funny in some ways. Right. Like, I think we were just talking few weeks ago about Gemini K3, we were like, who is running Gemini K3? Just given its like absolute size. And so there's kind of this very interesting thing where it feels like these kind of American companies that are in the open space are kind of playing for the smaller models.

18:58There's also an idea that, like, a lot of workload is going to be taken up on device versus in the data center. And again, I don't think it's either. It's not like a black and white thing, but curious about how you think about the division of like how things are evolving on what you're going to do on your device versus in, say, a big data center.

19:13The overall the models and the dense models will be more on smaller devices. Right? It's a small form factor thing. Models, a mixture of expert models have this like kind of demand that you have a lot of concurrent requests coming through because otherwise there are, you know, you're inefficient, like infrastructure wise, you're inefficient in many cases.

19:37You have to swap in weights because you just don't have enough memory to hold all these experts in, in memory. Right. And so I think we are going down a route of there will be on device or we will right size it. There will be on the device primarily because a you want privacy, b in certain cases you are just not connected.

19:59And then you have the larger models which are in the cloud and they are industrial scale hosted. That's how I think about this. Like if you have a mixture of expert model, you as an individual or as a small scale company, you just cannot afford putting the infrastructure down.

20:18So you will rent infrastructure. And so you will then, you know, go to some of these model providers and they will just host it for you. I think that if you look at the agent workload cases, there will be a very hard push. And this is what you're seeing, you know, in like internally with WatsonX, there will be a push towards specialized smaller models just to keep the cost in check.

20:39You just cannot afford the very, very large models. So we use them for advanced reasoning. And then we will have model routers which effectively make a decision where you go. I think the economic, the economics are in a way that if you are large scale enough, then you can actually specialize.

20:58But you may with this common infrastructure and this fluctuating workloads, you may have enough spare capacity anyway that you can just run the big model out. And so I think it's a balancing act of the model providers when they are choosing and the enterprises when they choose what.

21:20And I think this will be where these smaller models will play. Wherever you have the choice of actually bringing your per token cost into some reasonable amount, you will get a good, highly specialized or more specialized models. Or if the task is easier and the complex models will effectively split this out at some point to which model you could potentially go, just so that it's part of the integrated story.

21:42So I think this is how I see the market fall out. You cannot run these huge models on embedded devices, but certain stuff you want to run on embedded device and you don't want to move the data. So overall I think there is a there's a privacy argument and protocol in the enterprise that privacy argument's very strong just from the perspective of like regulatory requirements, like you just cannot move your data if you don't control the end to end flow of the bytes.

22:10Yeah, I think it will be kind of in some ways. I mean, the privacy point is something I want to pick up on. Rynne was basically that you normally we're like, oh, consumers are the ones who are really concerned about privacy here. But it always kind of feels like in the last decade people were like, yeah, they've sort of like loosened up a little bit.

22:24On that being something that like they're really, really hung up about. At the same time, it does feel like companies are super, super worried about, you know, particularly these AI companies being able to kind of train on their data and all this.

22:35Right. I guess maybe the question for you is Zuckerberg's messaging very much kind of focuses on the privacy benefits of having superintelligence. That's kind of personal. And on your device, do you think it really matters to consumers so much?

22:53Like, will that become a defining factor in what you decide to adopt in terms of an AI service? I think it matters. And I think the reason why we've seen that kind of change in terms of people seeming to care less about it has been, you know, the boiling frog scenario, right, where there hasn't been an option to get access to the services that you want to have without giving up a whole bunch of privacy, where, you know, you have to, you know, terms and conditions on everything you sign up for, where you're paying for it with your data.

23:22Right. So I think this is actually a really interesting for a certain type of consumer. And there will also be a certain type of consumer who does not care enough to set it up on their laptop or doesn't have a powerful enough machine, and there will continue to be hosted things for that.

23:41It's very much a balancing act where I think it is very nuanced here. And while yes, I personally care a lot about privacy, there's a limit to how much you can actually do about that these days, which is frustrating. Really a great discussion.

23:50More to come, I'm sure. I think we're probably going to have another set of Meta stories next week as they continue to kind of like gin the works and put more stuff out here, but I just think it's yeah, it's a really, really interesting kind of intersection of issues.

OpenAI’s Astra model

24:12I'm going to move on to our last topic of the day, and it's kind of a funny news story because it's almost not a news story at all. So OpenAI put out a blog post entitled Responding to the Next Frontier of Critical Cyber Capabilities. And Chris, it's kind of a funny story, because what they're sort of the substance of the post is we're working on a really great model called Astra.

24:33It's not released today because of concerns that we have. We'll be in touch. And so, I don't know, I mostly just wanted to flag it because I think it's kind of a funny story, and it does kind of feel like we're entering this really strange situation where, you know, frontier AI labs get a lot of juice out of announcing that they're not doing things.

24:50And what to make of that? I mean, I don't know, at this point it's kind of like, yeah, just let me know when Astra is available. I don't really want to know the details on the way. How did you respond? I guess to the story. I think if you had a large news story or your latest model in training hacked by a competitor's or a partner's system, and perhaps you did a whole YouTube video on it, which was in great detail on exactly how this worked, I think you probably would update your cybersecurity procedure for training models.

25:27I think that's a sensible thing to do. Yeah, you're actually sympathetic. You're like, actually, they're doing this for good reason. I actually think so. I think it was so unprecedented what happened. We knew it was coming at some point. We knew this day was going to come.

25:40But I think until it happens, you don't really realize the implications and therefore that forces you to go, well, actually, we had this procedure that existed before, you know, these things happened. And, you know, we've got to improve and we've got to hope that other people learn from these lessons and, you know, and can take the right precautions in the right way.

26:03And then actually, if you watch the kind of the Black Hat YouTube video, they made a really good point on this, which is like for the attackers. Do you know what I mean? This day has been coming. But he was they were basically saying, look, I don't think we have an example for large scale cybersecurity defenders where they're properly prepared in that way.

26:23And I think it was a really interesting point. So I think you have to update the procedures and you have to say, well, these are the things that we've learned. And then hopefully companies of the world will learn from that and then actually update their own procedures.

26:41Because I think the reality is as we move into cybersecurity, the speed and scale that it works at is really the key thing. And you have to sort of protect that. Now, you probably have to say there, well, you know, surely some of this stuff would have been obvious there.

26:55But hindsight's a great thing, right? Do you know what I mean? It's like, you know, I don't think anybody was expecting the models to be quite at this level of capability, including OpenAI. And, you know, and I think they've been massively transparent about it and really open and sharing everything that's happened.

27:11And, and I, you know, and I'm sure people will disagree with me, but I think it would have been easy for them not to be so transparent. So yeah, I sympathize and I'm glad that that they've been open about it. Volkmar I guess the question, you know, I guess Chris is bringing me around from making fun of them a little bit.

27:31But, you know, we do seem to have a real genuine cybersecurity issue on our hands. Maybe the other way to flip the question is, is Astra ever going to come out? Because it kind of feels like there's like two really big problems here. One of them is they don't seem to be able to really control the behavior of these systems.

27:44They're just really good at getting online. They're really good at compromising systems. The other bit is Chris, to your point, right. Like it's really hard to get defenders to up their security robustness. And so I guess your threshold is we need to wait until the world is safe enough to release this model.

27:59I almost have to take a look at that. And it's like, is that ever going to happen? Do they really have a handle on this issue or do you think it will be really kind of like solved in a substantive way in any reasonable amount of time? Yeah, that's a loaded question.

28:11So so I think like let's look at the players here. Right. So there is the cybersecurity hackers. There is the approval who are trying to protect themselves. And then there are governmental actors who, you know, have used this for warfare. And then you have industrial, the whole industrial espionage thing.

28:41Right. And so the question is who gets their hands on the model first? And so what you want to do is you want to keep your you want to keep your society safe, whatever that means. And that may mean that you want actually your cybersecurity department of your military to actually have access to the model.

28:56Right. And so I think it's a phasing. We are, you know, probably those AI lab vendors are working with the respective agencies, etc., and they're saying, okay, how do we phase it? And so that we get a competitive advantage for a while, you know, before all the gaps are closed without bringing down our society, knowing that, you know, adversaries are actually doing the same thing.

29:22And they also train these models and they also deploy these models. So from my perspective, we are just in the cybersecurity and cyber warfare. And so you work with the ones where you have the biggest exposure so that your, you know, your iOS, your Windows, Linux etc.

29:43and you're trying to close the holes, but you're trying not to close all the holes because you also have an advantage of not closing all holes. Right? And so I think this is a cat and mouse game. And we just offloaded a cat and mouse game from human brains, which have PhDs from MIT, into effectively the AI labs.

29:57And so the AI labs are all negotiating. So that's one that's one view on this. The other view on this is like, wow, what an amazing story. My stuff is so smart that it's dangerous for the world. Like that's The New York Times headline. Right.

30:15And I think it creates a lot of anticipation. And then sometimes, you know, it just becomes I mean, if you look at Claude, you know, it's like it feels like lobotomy is what they did to it. So yeah, it can for sure not do cybersecurity attacks anymore because it's so dumb that you actually don't want to use it for anything meaningful, right?

30:31Like, I mean, it's the fable, the fable five lobotomized. I mean, the model by now, in some cases is literally just unusable. I have definitely run into a number of things where it's like the behavior is definitely different and less good. Yeah.

30:46I guess, do you? I mean, to pick up on that, like there is a narrative, which is, hey, isn't this all marketing? And it is kind of funny. You don't normally see situations where companies are kind of like rewarded from a marketing standpoint for limiting the functionality of their model.

31:05But I also kind of believe that, I don't know, maybe our enterprise customers are really excited because they hear about this incredibly dangerous technology you can't release. I guess, how much credence do you give sort of more of the kind of marketing interpretation of this?

31:18Or do you think that there is really, genuinely some really serious risks that, you know, I think they're just trying to do their best to handle. So I think it's both. I do want to just remember that nine months ago we all wrote code by hand.

31:33That was nine months ago, right? Like the things have evolved so rapidly and so much faster than I think any of us anticipated. And that's very true. Where we now see that all of this stuff is moving faster and faster and faster and faster and faster, and that's both good.

31:46We can do a lot with these things. And also really scary in some ways too. And it's important to remember all of those things as kind of looking at the space in which this is operating. Right. Like many organizations can't get a deploy out in nine months.

32:05Right. Like, it can be very difficult for some companies to react to a security posture change that quickly. And I think that is the danger side. That is a fascinating place for us to be in. Right. And I do think that there is some aspect of, you know, we saw Claude get lobotomized because they overhyped it, right?

32:29Because they really hyped up the computer use class. And then they kind of had to from a regulatory perspective. And I also don't know how much of this is people trying to, you know, as a seated player, trying to get more regulation to preserve that moat as well.

32:44Chris, maybe we'll give you the last word here before we bring this episode to a close. I think one element of this that I keep coming back to in my mind is the open ecosystem, right? I think it's not anything. I think like the foundation, the frontier models are becoming more and more like, oh, well, we have to be careful about every new release.

33:04Meanwhile, open just keeps coming out with more and more and more and more, and the capabilities of even small, dense models as we talked about, are getting really, really good. How do you think about the security aspects of all that? Like do you think the open providers will have to themselves pretty soon, say, well, due to the cyber security issues we're also not releasing?

33:25Or do the incentives there just really kind of prevent them from taking those types of actions? Without quoting my new hero Mark Zuckerberg? It has to be open. I just I don't see how closed works. We need to put for superintelligence to be truly, you know, beneficial for everybody.

33:41It can't be concentrated into a few closed providers. That was the word. So there we go. But and I actually really agree with that. I think that the more the AI is shared across the world, the more that we share compute, the more that we're open with the models, the better chance that we are going to have to be able to defend.

34:02Right? And the reality is, if you've got an actor over here who has got access to these closed models because they're special. But then these people over here don't have access to the same class of models, then they're not going to be able to defend themselves, and they're not even going to know how to.

34:21Right. And I think we can't have this disparity. Right. And again, without open then the closed model prices are going to go up. We've seen that already. And you know everybody remember a few months ago we were talking about API token pricing and all that's disappeared because guess what.

34:40All the open models have become really capable again. And the closed model providers are going, oh wait a minute, these people are just going to leave. Right. And so we can't have that. We need to have an open ecosystem. It is going to be safer.

34:51It is going to be more secure. But you know what? It's it will figure it out. And I think Hugging Face was the perfect example. Right. They went to defend themselves. And what happened? It was like, oh no, none of the closed model providers are letting me do it because it's a cybersecurity problem.

35:10So they ended up reaching towards one of the Chinese open weight models. So therefore we need to open the playing field and let people be able to design their own solutions and do things. And I think that's going to be the case for even outside of cybersecurity, for superintelligence in general.

35:25If somebody has got access to great compute and somebody's got access to the best models, right, what happens when I don't know, you're in a legal scenario or what happens when you're trying to build a startup. You've got the best, you've got access to the best AI, but them over there don't.

35:39The disparity becomes the best way of solving this is open is always the best way forward. I think at some point, and I'm going to call it and I've called it before, there will be a point where open will overtake the closed models. That will happen at some point, right?

35:59It's just a matter of time. And when that happens, we're going to we're going to have some fun. And on that note, that's all the time that we have for today. Volkmar and Chris, this is a stellar panel. Always happy to have you on the show. And thanks for joining all you listeners.

36:10If you enjoyed what you heard, you can get us on Apple Podcasts, Spotify and podcast platforms everywhere, and we'll see you all next week on Mixture of Experts.