IBM’s mainframe chip collab, NVIDIA’s Poolside deal & Ox Alpha’s reveal
Visit Mixture of Experts podcast page to get more AI content → https://ibm.biz/~j8AlVLp98 On episode 122 of Mixture of Experts, hosts Tim Hwang and Aili McConnon are joined by Gabe Goodhart, Ash Minhas, Skyler Speakman to talk hot chips, mystery models, and major money moves.
Transcript
Chapters4
Intro
00:01this is a clear line in the sand from Nvidia that they are going to depend. Their future is going to depend on a strong, thriving many model ecosystem. All that. More on today's Mixture of Experts. I'm Tim Hwang and welcome to Mixture of Experts.
00:19Each week brings together a panel of brilliant technologists working at the frontiers of artificial intelligence to lead you through the week's news. On this week's episode, we've got Gabe Goodhart, chief architect, AI foundations technical content manager and AI engineer, and Skyler Speakman, senior research scientist at the Software Innovation Lab.
00:39Welcome to you all. I'm also joined by my co-host today, Aili McConnon, who's a staff writer at IBM Think. I always say this like there's a lot of news this week, but this week it felt like there was maybe like an overwhelming amount of news.
00:52So we're going to cover a little bit a new announcement about IBM's new next generation dual processor. We'll talk about Ox Alpha and the kind of mystique around it, which just as of today actually was revealed who's behind it. But I actually wanted to start by talking a little bit about what's happening with Nvidia, its recent line of acquisitions and and why it's doing what it's doing.
NVIDIA’s $6 billion Poolside deal
01:19the headline here is that Hugging Face announced that it was going to be acquired by Nvidia for $12.9 billion. This follows on sort of just another deal not too long ago, of a $6 billion kind of acquisition transaction with Poolside. And I guess, Gabe, you you flagged both of these stories for us because you were just like, this is all about open.
01:39And so maybe I'll kick it over to you to give your your hot take first. I think this is a clear line in the sand from Nvidia that they are going to depend. Their future is going to depend on a strong, thriving many model ecosystem. And right now that is open.
01:59And right now they see that that is a critical lifeline for their company. So they're willing to put money behind it. They've clearly started down that track with their work on the NemoTron model line. And the Poolside investment is clearly a further investment in making NemoTron a strong model line that people can rely on.
02:18And then this acquisition of Hugging Face is kind of just taking it to the max, right? Hugging Face is the cornerstone of the open AI ecosystem right now, and that's true of the models, but it's also true of the software. The Transformers package is owned and operated by Hugging Face.
02:41It is the reference architecture of literally every LLM on the planet that is released in open source. So every downstream architectural implementation follows from the Transformers one. If I try to get something merged in before it's merged in Hugging Face in Transformers, they say hold up, we can't count on this until that other PR is merged.
02:59So this is not I mean, this is clearly a play for the center of gravity, right? So Nvidia is making a clear statement here that they want to be the center of open source AI and that I think has some some pros and some cons. Right. So a big company like Nvidia that is clearly well moneyed, well respected and in a very strong position of power in this industry.
03:23It's great to have that support behind the open ecosystem. For those of us that live in that ecosystem, it means we've got backing for years, which is amazing. On the other hand, a big piece of the open ecosystem is enough that drives intelligence per weight downwards and in fact incentivizes making the model smaller and easier to run on a wide variety of hardware, and putting that behind a vendor that is clearly biased in what hardware model should run on, and the size and cost of that hardware creates a bit of a perverse incentive.
04:01So I'll be really curious to see how this plays out in terms of is it? You know, I think the obvious analogy that a lot of people are drawing is GitHub and Microsoft, where GitHub is the de facto standard for where code lives in the open. And thus far it seems like GitHub and Microsoft have done a pretty decent job of keeping a firewall between the incentives of the closed source Microsoft ecosystem and the open source code ecosystem.
04:25In GitHub, there's obviously product cross-pollination with Copilot, but they haven't detracted from the open source ethos of GitHub. So the hope here is that Nvidia would do the same with Hugging Face. And that Hugging Face would remain a shining light of open source championship and openness and that, you know, Nvidia would simply be putting their bet behind it.
04:50So we'll see what happens. I'd love to push more on what you've underlined. The sort of my question is who benefits more in this? Is it the open source community or Nvidia? You know what? In the best case scenario. How do they each help each other and what would be some of the concerns.
05:03And we can, you know, Skyler Ash, if you want to jump in to join the conversation, please. Yeah. Remember when Microsoft bought GitHub? It's basically that again. Now Microsoft has done very good good job of have acquisitions gone. I think that one of the interesting things with Nvidia and them acquiring Hugging Face is going to be that there's a lot of inference providers already providing inference to Hugging Face, and so how would that work now?
05:36Are they reselling the same hardware that Nvidia's providing them? There's going to be some interesting dynamics start to appear from that perspective. Open weights, open software, but closed hardware. Nvidia's play here is to get as many models trained and running on their CUDA, 'cause that's the larger draw here.
05:56So on the one hand, we do benefit from having easier access to a wide range of models with an asterisk as long as they run on Nvidia chips. So I think that's that's what's really going on here, both with the Hugging Face purchase and with the model factory from Poolside.
06:12They want to turn these models out as fast as possible, because they're still in the business of selling hardware. Well, I think on that note, it really leans into why this is important to Nvidia and why Nvidia is well positioned to be the one making this investment.
06:27Pretty much every other significant vendor in this space is in either the model space itself or the software space above the model. So if you look at any of the neo clouds that are providing that inference, if you're looking at any of the the big hyperscalers like their goal is to get you running, not it doesn't particularly matter what the chip in the computer is.
06:51It matters what the API surface you're hitting is and what metering gateway you're going through. Nvidia doesn't care. It cares about the chip under the hood, so there's no disincentive for Hugging Face to continue providing a wide range of inference providers.
07:05As long as all those inference providers are buying chips from Nvidia to power the inference that they're running through Hugging Face. So I do think there's an incentive alignment here that's fairly unique to Nvidia. Again, I think the big losers in this deal will be the companies trying to make commodity chips that are not Nvidia.
07:25Right. So we've got I think one of the articles I read about this clearly positioned this as a defensive play against all of the labs that are trying to do vertical integration and build their own chips, because if the big labs stop needing to buy Nvidia chips, Nvidia needs to sell their chips somewhere else.
07:40And so they're really going to aim at that commodity hardware like go broad rather than deep type of sales play. And the ones that really lose out are the others attempting to sell broadly to, you know, general data center consumers. Yeah. Can I play bad guy here for a second?
07:56Because obviously there's some discomfort in the room with, you know, Nvidia's nefarious plan here to keep everybody on their kind of closed hardware. You know, isn't there another view here which is like get real, right. Like basically these commodity players were never, ever going to catch up.
08:09Nvidia was so far ahead. We're just recognizing reality, which is that like the game has been lost in terms of certainly open hardware, but maybe even commodity hardware here, just given Nvidia's like extensive lead. And if any do people agree with that.
08:26Like this is in some ways realism versus them really kind of closing off an avenue for commodity catch think to me. Yes. And I also think it's important to note that the smaller hardware vendors that are trying to do commodity hardware that are not Nvidia, they're trying to soak up excess demand, frankly, like right now we are in a very skewed supply demand ecosystem.
08:48And so this isn't going to change that dynamic. Now if the demand really does drop off, that's where you're going to start to see this become a bit more of a food fight. And you're absolutely right. You know, Nvidia is going to throw its weight around.
08:55And, you know, maybe it's just calling a spade a spade. But I think right now all of the possibility of other vendors losing out is just speculation. Because frankly, if a data center needs chips, they're probably going to look for Nvidia first.
09:14But if it's a, you know, nine month waiting period, they might look elsewhere. And Nvidia is probably okay with that because they're still selling as many chips as they can I think it's also interesting whether there's like a new, you know, taking this news.
09:25You mentioned, obviously, the licensing of the Poolside model factory. Is this like a new, bigger, broader playbook for AI consolidation? Because like in that case, what struck me was, you know, they weren't acquiring it. They, you know, bought the license.
09:39They're hiring the team, you know, so sort of Nvidia kind of grabbing land in a variety of different kind of manners. But perhaps that's too simplistic. I think a key part of that Poolside Nvidia conversation was this non-exclusive license.
09:56So Nvidia is not necessarily doing a land grab there. They are going to be one of perhaps multiple companies that are going to be using Poolside's approach to model creation. So that that was kind of in the details of that agreement. So yeah, we'll be interesting to see who else might be in line to pay $6 billion to Poolside for their license.
10:15I don't think there'll be too many customers left in that. But this is not exclusive Poolside and Nvidia in that particular arrangement. Yeah, that's maybe where I want to kind of like do the final few questions on this segment is, you know, normally when you spend 19, $20 billion, you're usually a little bit short on the money.
10:29But we're also talking about Nvidia here where they could go for a lot more. I guess the question is if I'm curious if you have any thoughts on like who's who's next up here, where else would you go if you were trying to kind of consolidate in the open space?
10:44Do you think Nvidia is going to kind of make another another play, or is this kind of like it for a little When we look at sort of like what what does the AI world look like? I mean, the stack is like, you know, chips and infrastructure. Then it's like the cloud providers providing that, and then you've got the application layer.
11:03I think with both of these things, it's like infrastructure locked down. You know, they're in that middle layer now. And like they want to make really good models. They're going to use obviously that Poolside acquisition to make their own models better.
11:21I'm curious. And I mean, I know I'm responding to your question, Tim, with a question. But the question is, is like, are they going to go even higher up the stack and start looking at some of that application layer and maybe going for a slice of that somewhere and make in making that open or, you know, is it, is it like, you know, going to go grab Open WebUI?
11:35they've already taken a small step in that direction with NeMo Claw. Right. So they have put their probably one of the largest companies behind an open source agent harness at this point. So you know that there was a big announcement a little while ago about NeMo Claw.
11:50And it's kind of I wouldn't say it's it's petered out, but there hasn't been as much hype about the Claws as of late. I think a few other agent harnesses like Hermes have taken up a little bit more of the oxygen for the generalist agent, but, you know, I guess I would be surprised to see Nvidia try to go the actual like for sale buy this as a consumer software platform route.
12:26You know, I don't think we'll be going to like AI and having our agent chats there the way folks do with ChatGPT or Claude, it just doesn't seem like that incentivizes the right things, because ultimately their business is still built completely around the chips that are under the hood.
12:39So I think for me, the two recent deals we're talking about here are just a clear sign that they are going for as many flowers blooming as possible, right? And if they can put their money and their mouth behind specific projects that they think amplify that goal, they will.
13:00But I don't see them trying to then also carve off a corner of the market and say, but forget all those other flowers. Come to our specific flower.
IBM’s new mainframe processor
13:11I'm going to move us on to our next topic, which I guess is related in some ways, but it's a it's a different side of the market entirely. We've talked about the mainframe market here on MoE in the past, and one of the things I love about the market is that it's very different from what we know of usual, you know, things around AI where the general approach of AI industry has been just like, all right, let's just launch it and see what happens with mainframes.
13:32You really kind of can't afford that in some ways. And I know you flag this one Ash if you want to introduce it, but it sounds like IBM is basically out with a new dual processor in the mainframe space, which is a kind of big deal just given how cautious people are in mainframes.
13:49Do you want to talk more about Sure, sure. And yeah, they presented this at this week was Hot Chips, which is a big kind of semiconductor annual conference kind of chipmakers. Researchers come together, unveil new processors and designs. And, you know, many of the participants were some of those games that you mentioned, the frontier companies who've all come up with their own kind of custom chips that they're they're sharing.
14:09But in this context, IBM was announcing kind of a different type of chip, you know, this dual process architecture. And so this one they've developed with Arm, the AI software ecosystem. And it was different about it is it combines IBM's Z.
14:25So the mainframe workload with Arm and all on one chip and with the ability for the ultimately to be able to switch between those two workloads in nanoseconds. So I think the bigger picture thinking is that there's a growing challenge in enterprise AI, where organizations want to run their inference closer to where that data resides.
14:48And the reality is that much of that data sits on mainframes. You know, many of those IBM Z and Linux one systems. And then at the same time, you have this Arm ecosystem where large share of the modern AI software is being built. So I think that, you know, the idea is rather than moving the data between these two places, IBM is sort of proposing this chip to bring those workloads closer together.
15:16It may be, Skyler, if you're game to jump in on this first, I'm interested, you know, from an infrastructure perspective. Why does this matter? How big of a deal is this? If you get all the way down to the nitty gritty details about how these processors work, there's two different philosophies.
15:29One of them is doing multiple simpler instructions. So if you want to add two numbers, there's one instruction to go read a number from memory. There's another instruction to add it, and there's another instruction to write it. The other philosophy is a single, more complex instruction that does that all end to end.
15:54And so computer science over decades has evolved in kind of those two different philosophies. And it's really interesting now to see these worlds somewhat collide in the IBM Arm space. I do not know the technology behind it about how they're doing both of those instruction sets on the same chip.
16:07Kudos to them. Very smart people working on that. But it really, I think, is going to be interesting to, you know, ask questions like, is your bank going to go to the App Store and update its core banking software? You know, can you go to an ATM and do a transaction and say, wait a minute, our back end is training a model right now.
16:27Give us, give us 20s. And those are questions we had to ask before, because these things have been completely separate ecosystems. And now. IBM and Arm are saying still mainframe functionality, but the ability to run these two different instruction sets and they have again, what I want to hear is two different philosophies of how you go about writing code, coming together under one place.
16:50And the thing that it kept bothering me when I was reading this is, on the one hand, it's cool to see these fences being brought down, but there's also this approach where you should be asking, why were those fences there to begin with? Do we really want to have Arm software sitting in the same places as, for example, our government services, banking, banking services and other things that have relied on mainframes for decades?
17:14So watch this space. I don't know who else else can be can comment on that. But this is I think this is bigger than just a single chip sharing two different instruction sets. It really is two different worlds colliding. And it'll be really cool to see it play out in the coming months.
17:29And I think it sort of yeah, I think to your point of it being a larger question to it, this looking at enterprises trying to integrate, you know, AI with these mission critical systems, these established systems that, you know, they can't just totally rip out and then kind of bring in something totally new.
17:49And the analogy you use of the App Store somewhere, and as I was talking about this news that explained it like that, like suddenly, you know, it's like mainframe consumers, you know, have this whole app store of options, you know, if they're an enterprise, you know, hoping to bring in the inference kind of closer to where their data is sitting.
18:10And it's harder to get out of. I think both IBM and Arm are winners in this case as of now, I think I think right now I think this really is a I don't think this is a zero sum game here. I think there's opportunities for for both of these kind of established players.
18:25And it's cool to see an announcement back in April of the partnership. And then a few months later, the announcement of the chip. We'll still have to wait for its actual release. Maybe maybe I'm a bit doubtful for that, but it's cool to see that sort of quick turnaround on this type of So I want to bring in the perspective of a software engineer on this, which is this is going to make a lot of people's lives a whole lot easier.
18:43So, you know, a few months ago, there was a big kerfuffle in the market that certainly affected us at IBM, where Anthropic announced their migration tool off of COBOL. And the market worried that that meant the death of the mainframe. And we had a lot of well-articulated responses out of IBM.
19:06But I think the core of all those responses was, look, COBOL is a means to an end, but the end has not gone away, right? The reason mainframes exist is bulletproof reliability. Like I was talking to, I forget who it was earlier this week, who was literally part of the team that said you could literally take the mainframe out back and put a bullet through it and it would not stop processing transactions.
19:28And that's just not true of any other piece of hardware, right? Like, you can you could shoot a bullet through a RAM stick on the Z system and it wouldn't blink. Now that reliability is what the foundation of many of our most critical pieces of infrastructure is built on.
19:48And so, yes, there's a lot of code trapped in COBOL, which is a pain in the neck and is probably very difficult for companies to find programmers that can actually handle that code. But the need for that reliability has gone nowhere. And so this is an attempt to say, well, we can solve that problem in a different way, rather than bringing the mainframe code out and then sacrificing that reliability that you've come to depend on.
20:09Why don't we bring the modern code in? Right. And as someone who has tried to cross compile code for arbitrary random architectures, if you're thinking about import torch as a Python program and you're like, what does cross compiling mean? Well, let me tell you, Python is a C program.
20:31Every single library that you import in Python is in fact delegating down to a C library if it's got any kind of performance behind it. Torch itself has backends compiled against every single accelerator architecture out there, so it is a massive pain in the neck of, you know, ecosystem, targeted cross compilation.
20:48And the idea before this announcement of taking something as complicated as an LLM that is backed by all of this complicated acceleration logic and cross compiling it to a completely bespoke, bespoke is the wrong word. Completely firewall and separate chip architecture, both for the accelerator and for the standard CPU processing, was a daunting task and this was the the purview of, you know, potentially years of work to get, you know, individual LLM architectures ported over to a Z system.
21:20Now with Arm. And the fact that a huge number of people are running Arm workloads on their laptops, on their phones, even starting to be on their desktop processors and server processors, means that a huge amount of software has already been adapted to the Arm ecosystem.
21:41So that work is done for us. So this is hopefully going to open up the floodgates of bringing a ton of interesting workloads to the mainframe that just could not run there in any reasonable amount of effort before. Now to your point, Skyler, it'll be interesting to see what that does for the reliability of the other workloads that are sharing the processor.
22:02So the proof will be in the pudding there, but at least on the surface of it, this looks like this is going to make a lot of people's lives a lot easier. should we talk a little bit about the future? Like I'm thinking about like in 2050. You know, will we still have mainframes?
22:15It feels like this technology that people like really gripe about and are always like, the mainframe is on its way out, on its way out. But to Gabe's point, right. Like it is, it is rare to find reliability of this kind. And, you know, I guess the question is, I guess in some ways you could almost read the trend line.
22:30It's like, actually, mainframes will be here and sort of bigger than ever. But do you buy that? Because I'm kind of curious about whether or not this, like pretty dusty kind of ecosystem in some ways is like much, much more robust and interesting than it I think that there are certain workloads where latency really makes a difference.
22:50Okay. And the latency. Issue just doesn't really exist in mainframe. You know, you're getting like, transactions processed in such a fast time that we will always need the ability to do that with certain types of transactions. And so I think that we will trend towards sort of bigger, more powerful computers of this kind pretty much for ever, really.
23:17The, the, the point that Skyler and both gave me. I also like sort of have some questions I would say right now and curiosity. You know, this is a great sort of thing that's come out of this collaboration with Arm. And, you know, I guess the when I think about this from sort of like a computer architecture perspective, they could have just put some Arm cores into the chip and they didn't.
23:42Okay. It's all combined in one. And that's really, really cool and really, really interesting. I would love to like, go and talk to the people who made that architectural decision and go, why did you decide to go the hard way? Because there could have been easier if you did it the other way.
23:55And yeah, also that then raises questions of like, well, you know, we're trying to process these things on on mainframe today, which are like sort of millisecond type transactions. And so now if you've got, you know, I don't know, let's say PyTorch running on there and something else.
24:08Okay. What does that do to the behavior of the overall system. And it's that going to mean that it will like it will mean a much, much bigger computer in 2050. yeah, that would be a really funny outcome is just like we're living in 2050 and it's like giant mainframes that that'd be really interesting to see.
Ox Alpha, Z.ai’s newest model
24:32I'm going to move this on to our kind of last topic of the day. You know, I think it's an adage in 2025 that it feels increasingly like we're living in a cyberpunk novel of some kind. This story, I think, was very much like it basically on OpenRouter, there's this mystery model called Ox Alpha, that kind of hit, and people were very impressed by its capability.
24:49And most interestingly, at least in the announcement, it was claimed that it was able to serve 100 trillion tokens a day. And you know, which immediately raised questions about like, what kind of compute are you running to be able to offer this kind of thing?
25:03I think as of this morning, it has finally been revealed that it is GLM-5.3 open source Chinese model. And I guess maybe the first thing is we should just do the usual vibe check test thing is we Gabe, Sky, I'm curious if any of you kind of played with the model.
25:19What do you think about it? Just from a you know, what kind of like a wine review? Like, what do you think about it? Is the hype justified? Benchmarks are one of these things that I think in this like probabilistic world are going to just be contested forever.
25:32So at some point, I don't know, about a year ago I just put together sort of like, you know, my own little like Benchmark suite, whatever. And I ran that through OpenRouter with it. And I thought, you know what? This is actually pretty good.
25:44Pretty good. I mean, it's great that it's free. It's not free now, but it's still dirt cheap. And I was pretty, pretty impressed with with the output that I Yeah I mean I gave it the smell test as well. It smells like a good model. But like with wine, I at this point can't actually claim a connoisseur's palate because most of the workloads I want to run can be satisfied by a sub 30 billion parameter model on And thank goodness for that.
26:13So you know little little plug for small models there. But you know my my feeling is that it is fantastic to have competition at the top end. Love seeing it. Very curious whether this throughput number that they are claiming is due to breadth of scaling and just raw compute, or whether they've done something truly novel in the model architecture that enables massive throughput improvements while maintaining this top end quality.
26:38That would be genuinely very cool if they figured some tricks out about how to, you know, get the bits through the attention mechanism faster. But at the end of the day, we have a glut of very, very, very good models. And that's awesome. And I'm still going to try to run as much as I can off my local machine.
26:57My interest in this story is as much about how they decide to go secret. And, you know, I'm interested in your thoughts on, you know, presumably intentional. Does it build hype? You know, is there a more to it than that, that they, you know, it was going to be revealed?
27:12You know, at some point, you know, why the why the stealth mode. I think the stealth mode worked. I mean, in between the time where we had these topics chosen for this, for this podcast and when it's actually recorded, it came out as GLM 5.3.
27:24Would we be talking about the release of GLM 5.3 if that was just how it came about? Is it so much more fascinating to have this idea of a hidden model? As Tim pointed out, it's like a cyberpunk novel where you can have these, you know, mysterious people showing up and competing in a, I don't know, a medieval joust with a with helmets on it.
27:45We don't know who they are. And so I think definitely well done to the makers of the model to release it in this way. They got the hype they wanted to a few days of great speculation and then the release saying it is just an improvement, not just it is an impressive improvement over one of their previous models.
28:07So I think, I think they really did a great job with the the hidden reveal, letting let them the internet talk about it for 2 or 3 I will say another thing to that, which is that timing a model release is very hard. You know, we at IBM just released the Granite 4.2 models, which, you know, we are proud of.
28:29They are certainly not going to win the benchmark race, but we think that they're going to be very useful for the target audience. However, a day after we did that, we had Qwen drop yet another benchmark busting, amazing model that also runs on the same local hardware.
28:45And, you know, picking when you want to release and how you want to release is a a very difficult game of speculation on who else is going to be releasing competitive models in the same size, in the same space. So, so like you said, Skyler, I mean, kudos to the team for choosing a different route to release, because if this had been just another amazing Chinese model, like what a world we live in, by the way, you know, would we even bother to talk about it?
29:09But choosing a release strategy that almost was competition proof because nobody else was doing it? No, I don't think anyone else can probably pull this off again for a little while. I mean, we had it kind of with nano banana for a while. This isn't the first time a lab has sort of stealth launched a model with a catchy name, but it is a good strategy to mitigate against.
29:37Oh, shoot is one of the other frontier labs also going to release their new model on the same day or the same week? Who knows? Or sometimes within, you know, I'm recalling one OpenAI Anthropic pair of releases within like an hour of each other where it's sort of it becomes a sort of playground battle of who gets the most attention.
29:56I think that they did have like Like an announcement to say, oh, they wanted to release it in a sort of stealth mode so they could get unbiased, actual, real feedback from developers to to know, you know, how good their model is. In all of those use cases.
30:10I do think there's probably an element of that in like why it was stealth, right? You know, they didn't want to say, hey, it's a Chinese model. And for it to like introduce any sort of bias into determination as to the models performance. I do also think that there's probably 50% of it is marketing, right?
30:30It's a great marketing thing to That's great. Well, that's all the time that we have for today. Gabe. Sky ash was great to have you on the show, and Aili is great co-hosting with you Thank you. thanks for joining all your listeners. If you enjoyed what you heard, you can get us on Apple Podcasts, Spotify and podcast platforms everywhere, and we'll see you all next week on Mixture of Experts.