Deepseek’s Impact on Your AI Strategy
In this episode, Keith Townsend is joined by Dion Hinchcliffe—VP and Practice Lead for The Futurum Group’s CIO practice—to explore the monumental impact of DeepSeek, a new large language model (LLM) that’s challenging AI industry norms. They discuss the astonishing claims surrounding DeepSeek’s minimal training costs, its potential to run on everyday hardware at ChatGPT-level [...]
Transcript
All right, it's been a while since we've published a podcast because we're trying to only say stuff when there's something to say and man, oh, man, what a week. There is absolutely something to say. We're recording this the week after Deep Seek has dropped. And I have no better guest to talk about this. Couldn't possibly be a better guest than Diane Henscliff, who is our VP and practice lead for our CIO practice here at the Futurum Group. Diane, I think it's the first time you're on the program, right?
First or second, I can't remember. Yeah, something like that. Well, the the the details of which I don't know matters, but I wanted to bring you on because you had a really great, informative post on social breaking down Deep Seek. Now, I wanted to share with our mainly infrastructure focused audience why Deep Seek matters, et cetera, et cetera. But let's start out there. What is Deep Seek and why does it matter? I. I. and Anthropic and Google and all of those folks.
It is a is a large language model. It is from a company in China. It is a very high performing model. It's the highest performing model that's ever come out of China. They've been producing some LLMs. But when you look at the benchmarks, there is now very well defined benchmarks that allow people to to quantify how good is this new new language model. There's all these benchmarks around reasoning and other things that you make them go through. And Deep Seek landed at number three right out of the gate, which it's never happens.
I. across a wide swath of metrics and number one in things like coding and math. And that wasn't a big enough deal. It reportedly took them forty five times less training effort to create this model. So it's a dramatic reduction in cost and time, which is the big input for these models. That's the moat around the whole industry is that you have to come up with hundreds of millions of dollars to train your model. And they somehow haven't had to do that.
And if that wasn't enough, it could run on a regular computer. I mean, a regular high end. You better have a Mac studio or something. We can get, you know, oh, one performance as opening eyes flagship model on your own, on a personal computer. I. and their ilk. So it's very disruptive. It cost a trillion dollars in market cap to vanish from the tech markets yesterday and videos down 17 percent. Look at the stock prices, the cliff, because if they're if they're if they're right, if they're accurate and it's not true that they are, there's a lot of speculation that they have 50000 H100s in their back pocket that they actually use.
But they can't pay that due to expert controls. I. Yeah. So there's a lot to unpack there is a Chinese company, as you mentioned. And the claim is that they did this with basically six million dollars worth of infrastructure. That claim might be a little bit dubious. Exactly. Even the case, let's say let's accept that is, you know, 50000 H100s. One of the things that I wanted to help with context is H100 performance versus the latest NVIDIA chipset, which is I think the B200.
I think it's fair to say that B200 was like four times more efficient than the H100. Mm hmm. Yep. And higher bandwidth. Yeah. Right. Yeah. So higher bandwidth, more, more efficient. And the. Implication of that is twofold, one, that this company has done some amazing things are just unquestionably the second question is. What does that mean downstream? Because I don't think most of us are going to develop models. In most cases, I think that that's still going to be, you know, a fairly high end request, regardless of the barrier of cost, there's other barriers such as talent.
Time, et cetera, and access to data. What does this mean for the downstream for the enterprise? So for the enterprise, it might mean that a world class is now going to be repriced completely differently. Now, the big risk is and a joke is that they really should call it chat CCP, right? That this is this is run in communist China. This is hosted there if you're using the online version of it. And it does seem to store keystrokes and and your all your history and all your output in on China's servers.
But you can it's open source supposedly to ever try to determine how open source it really is. And you can put it on your own computers. So technically, if an enterprise wanted to get access to absolutely leading class output, that would cost a fortune through through your open subscription, enterprise subscription. You can now do this for for essentially free in your own data center. And so if you trust the source and there's been a lot of questions, it is censored. But for example, you know, you can't ask it about Tiananmen Square.
It says it won't do it, but you could ask it about anything you want about America politics and it will answer. So it's tuned very carefully to CCP sensibilities. And while a lot of people have jailbroken that and said and discovered the information is in the model, it's obviously clearly censored. And it was very interesting. If you're willing to take the geopolitical risk of sourcing leading class AI from China, then the price. I mean, and the price point makes it very attractive.
You know, one level model can do remarkable things, produce vast amounts of very high quality output, research, develop software for you. Not a level that if you've just been using chat GPT, you don't understand what a frontier model can do. It can produce, you know, ready working applications with almost no defects that are relatively sophisticated based on the spec. At the push of a button. And that's what you'll be able to get with DeepSeek in R1. That's the specific model.
They released another one, by the way. I don't know if you caught that, Keith. They have another model. They just released a multimodal model. That's really more focused on analysis of images, text and production of images and text. But anyway, it's everyone's worried that this is going to make everyone's capital plans this year absolutely upside down. But it's great for enterprises if it makes the price of AI ten times cheaper, which is likely to do so. R1 has some very amazing optimizations around.
It only uses eight bit floating point. It uses extreme compression on the keywords so that as some like an 87 percent savings in VRAM. That those are big deals. And those are those will get to every model. Everyone's going to copy their optimizations for sure, at least because it's open source. They can see how we how they got it onto a massive model, onto a small machine. Everyone's going to do that. I think it's going to be great for AI on the conspire side.
The bottom line. Yeah, I was talking to the chief scientist over at NVIDIA. This was, I think, a GTC last year. He was telling me that they were actually going in the opposite direction, going to smaller bits or to get to deal with chonking and all of these challenges with it. And if you read the paper and I'll link it in the show notes, if you want to go deep and read the research paper from the deep sea folks, it's amazing levels of.
Ingenuity to optimize this reinforced learning method, go back and watch 100 days of AI to learn what reinforced learning is versus supervised learning, et cetera. But they are standing on the shoulder of giants. And I think as a result, the enterprise will be able to stand on the shoulder of deep. I just saw a analysis. Someone had three in Mac, Mac, Mac, Mac minis or Mac studios. Daisy chained together running deep at the same performance level of a chat GP or one.
So this gives the enterprise an idea of what this means for you. If you're looking at Emerson projects. So, Diane, let's put our crystal ball on and help customers really digest what this should mean for their capital expenses. If you're looking at building a solutions over the next 12 to 18 months. How should you begin to rethink your capital investment in big GPUs, et cetera, or should you rethink it at all? I think there's some obvious things you should not do.
First of all, one of those is don't make any long term. Don't sign a long term contracts. We'll see what happens. At the very least, inference has been made dramatically cheaper. Now, the question is, is training been made dramatically cheaper? And that's going to have big implications. Because, you know, we're trying to we're trying to beat the market and create a leadership and AI in the United States by launching massive capex projects like Project Stargate. You know, 500 billion dollars to create models and no one else can do because they can't train them up.
They don't have a training environment big enough to go that deep into the data. This could set all that this could upset the Apple car completely if they have somehow created a 45X breakthrough. That's basically almost two orders of magnitude breakthrough and training efficiency. I don't think that's going to hold up on the training side. On the inference side, it's obvious because everyone's able to replicate those results. So I would say don't make any long term plans. Don't make any big commitments to budget or vendors, because I think it's all going to be different.
And I think it's just going to be an arms race. It's just everyone's going to get table stakes and we'll be able to reach the next level of artificial intelligence, something I call super intelligence. It's not AGI, but this is going to allow us to run models that have super high capabilities and can do things humans can't do. D. human level, they can't do anything that humans can't do yet. That threshold is likely to be crossed this year, if not for sure next year.
And these types of efficiencies are going to allow that. So you can get these super powerful artificial intelligence that can solve problems that humans and enterprises haven't been able to solve up to this point. That is what Sam Altman is really after. And that's what Anthropics is really after, is AIs that can do things that no one else can. And they're the only ones that own the AIs that can do that. I don't think they're going to be able to retain control like that.
I think their dreams are not going to come true. So don't get locked into those vendors. Keep your options open. The AI space is going to continue to have these kind of dramatic breakthroughs. This is a chat GPT level breakthrough, the first one we've seen since chat GPT. I think, you know, we're waiting for the next big one is going to be the one that has an AI that can is a super intelligence that can do things humans can't. And it's going to probably be down this ultra hyper efficiency path.
So let's end with a plug for some of the work you've been doing on the future and group side. And this is adjacent to this conversation. You came out with a big CEO data report around just the success of AI projects. This is a partnership with Kearney. What are some of the adjacent learnings and how will people find out more about this? So, yeah, we just launched a landmark study on what CEOs are doing with AI going into 2025, what their short and long term plans are, where they worry about with it.
What are their challenges in rolling it out? And we discovered that 59 percent of CEOs are leading AI, which is a surprising stat for us. But we're also finding the ones that micromanage it or over involve themselves in it. They report poorer results than the ones that step back after they set the mandate. Everyone must take part in transforming, doing AI transformation on their part of the organization and then stepping back and making sure they have what they need. Those CEOs that do that, that take that step back and let everyone in the organization innovate are reporting higher success rate.
That said, we did a lot of in-depth interviews with CEOs talking about what they're doing. It is so top of mind and they're looking, they're going after cost savings right now, better customer service. And they had great, very detailed stories about how you can go in now and just ask a chatbot, you know, how many unshipped orders do I have? How many credits do I have left? Do I, you know, what problems do I have with my shipments in this part of the world?
You just go and get these answers. And customers love it, not having to sort through reports and long dashboards. They can just ask the question and get the answer. And so it's a crawl, walk, run thing. They also keep, have an eye on being strategic. You know, what are we going to do to reinvent our business? They think AI, for the most part, not all of them. If you're in manufacturing, they really only think it's going to affect operations and supply chain.
But many of them want AI to rethink and create unbeatable new products and services, new business models that will generate new revenue. They think AI can do it. They think only they can do it today, but they're preparing for that. So it's very exciting. If you want to see it, you have to request a copy. com. You'll find our website and you can request a download. We'll get a copy to you. Diane, I appreciate you taking out time on your busy schedule.
You're in demand. People are DMing you, asking you to appear on shows and explain this deep seek disruption to everyone. I appreciate you explaining it to the audience. Links to Diane's socials will be in the show notes. com. Diane also published just a couple of months ago. He's been a busy man, a CIO data set with equally interesting insights. You want to find out more about the CTO Advisor. com. Talk to you next CTO Advisor podcast.
Thanks, Diane. Thanks, Keith.