Why Enterprises Need a “Reasoning Layer” for AI Apps | CTO Advisor Lightboard with Articul8
Transcript
All right, we're still on our CTO advisor road trip. We're heading into GTC next week Arun. And Arun, we had you on 8 weeks ago. And this just goes to show how fast AI is moving. We talked about the whole no apps concept and I'm going to challenge you on that. Uh, my good friend Patrick Moorhead, who is a proper CEO level analyst is building apps with claw code. We have open claw, etc. So, I'm going to push back and tell you we have way more apps than we've had before.
So Keith, first of all, thank you so much for having me again. It's wonderful to be here. Actually, you're proving my point. You're talking about Pat Moorhead, but uh, let me tell you a story about the program managers at Articulate. Okay? In the last uh, I would say 8 weeks or so or slightly longer than that, we have our program managers push about the same amount of production code into production applications than our engineers have. Okay? And the reason for that is these are program managers who understand what is the outcome they need and they're building apps for themselves, not for anybody else.
And when I said there are no apps, what I meant by that is there are no traditional apps where you have a a product manager defining what the app should be an engineering team taking multiple sprints to build one and you have all these teams to actually manage this app and maintain it. So, all right, you're breaking my enterprise IT head. When you say they're pushing code or applications to production, define that for me. Because when I think about production applications, I think about the stuff that your team builds for customers.
So, there's two ways to think about it. And it turns out first of all, you have apps. Apps, but I'll put a star on them. Okay? This is what apps people are building for themselves as well as scaling those apps. Going in and once they built it for themselves, they can get other people to use it as well. But those can only be built inside an enterprise for usage if you have a what I would call a platform that supports building it safely and easily.
So, it's a software development kit from the platform you're using and all the heavy work that goes into making sure that these kinds of things can scale, they're secure, they're safe inside an enterprise is under the surface. What they're building with the new tools, even the tools that came out in the last 8 weeks or so, can use this to be safe. However, they don't really need to know anything other than somebody telling them, "Look, this is the box that you're given. " All right, so I hear you and I I still need to connect some dots because they're producing put pushing production code that solves an internal business case.
You're like, "You know what? " They can now build that. They We talked about this reduced friction. But I think about like, "Okay, now I've created code. There needs to be code reviews. " And typically, that's complexity that I born as a IT enterprise expert. So, that is part of the SDK along with the build tools. Exactly what you mentioned, which is code reviews, uh, pull requests, making sure that the safety and security responses are done. You're doing not just uh, like say a surface level security scan, you're actually doing defense of depth kind of attacks and then figuring out is this all still safe?
" it will get it done by circumventing the rules. So, you need to have these two combined together to get you the apps. But what I'm telling you is even though these tools are maybe 8 weeks old, maybe 9 weeks old in the world, they're at a stage where you can still have only one human check here and push to production. So, I'm getting this. >> Yes. My I I get the concept, but in reality, what happens when I create more code?
There's more cost. Let's talk about the economics of what it means to support that. >> Perfect. So, it's not just more code, but also in terms of what is the turn that you're going to experience, right? So, because when you go from a world where maybe your team is able to support 10 production apps to 1,000 production apps, how do you actually do that? And the apps also no no longer required to be say alive for 2 years or 3 years or 10 years.
They may show up for a couple of months and disappear. But in the couple of months, they need to still be Yeah, they they they need to be secure, they need to be stable. Uh, there's bug fixes, there's uh, there's always patching. Like the what version of JVM blah blah blah does this this this Yes. So, all of that is what I would call the actual iceberg and then you're only talking about the tip of the iceberg, right? So, if you're talking about the platform and below the platform, in any of these kinds of tools, of course, you need to make sure that you have the right model access.
But even if you're using the models in production, say in the cloud, any more any of these models you use, they have a variety of cost like say models in there, right? So, you can do auto and be okay more or less because it doesn't go to the highest thinking mode ever. But if you go if you turn on high thinking, for example, in one day you can blow your budget. " But why would I not choose it otherwise? Yeah, so most people like they wouldn't realize the difference between auto and high thinking and would think that well, if I use high thinking, I'll always get high quality.
Not necessarily the case. It depends on what problem you're trying to solve. But that's just again a very small portion of And we call this commonly between Articulate and the CTO advisor, we we call this the reasoning layer. Someone or something needs to make this decision. The human needs to decide what business outcome and what process they want. And then the system needs to meet these levels of decisions. That doesn't Arun, sorry. That doesn't exist. Not in production.
That's exactly what we're building and actually delivering to our customers. But if you're looking at cost, right? So, this is just if you're going after a model that's outside. What if you're an enterprise that's deploying internally? You need to worry about GPU cost, of course. But then most people stop here. But below that, there is an entire cost module that is not really that well understood, which is you have to think about, "Okay, you do all of this. " How are you going to deploy all of this and make sure that you can do this?
Now, I can tell you even in 1 week, five models land that are at frontier level or above frontier level, there is no argument that you have to actually update. Right? It's not like you can choose not to update these. Yeah, this is no this is no longer, you know, do I go to the next version of SAP to get features I don't need? Users, we talked about this preamble is which is, you know what? Users are abandoning their internal co-pilot systems and using what they have available openly commercially.
You know, I'll I'll whip out the credit card and it's nothing to spend $200 a month on the latest clawed or open AI. Really doesn't matter. >> You want a GLM 5, you want deep seek, any Like users are I've seen people spend their own personal money to get these things. Because the the delta that they get in terms of performance and quality is something they can't beat, right? So, that's why you have to be prepared for this. So, there's a deployment infrastructure that you need to be able to do.
How do you version it? How do you make sure that 3 days ago I gave an answer. Now I change the model, the answer is slightly different. How do you make sure that the users understand that? That's number one. Second thing is you need an observability stack. Hm. Who's using what? By how much? Which outcomes are coming out of what? All of that needs to come from an observability stack, right? And which by the way, today most of enterprise teams would build themselves internally.
They will have to figure out how to partner efficiently because this level of change nobody's used to. And not only that, this level of data we're not used to. Not used to. Not at all. Like and like and then there is a security stack. And you're talking about all of this, okay, then how do you then go do uh, all of your updates? Not just at the model level, but then once you change the model, people go and change a whole bunch of prompts.
And then if you have your own personal prompts, and you have your enterprise-level prompts, how do you manage all this? Right? And if you notice, this is soon becoming a massive massive iceberg. And most of the time, the conversation starts stops here. Maybe it goes down up to here. And this is the problem that I've been talking about for the past few months, which is how do I scale the 10x developer? And the 10x developer is not even a engineer anymore.
It is a program manager, product manager. It is the person that has kind of figured out the heck, but I don't know how to make that scalable across my entire organization. >> Exactly. And that's I mean, I've not even written a couple of other things, right? So, for example, the other thing you need to worry about is what is the SLA you need? Do you need the answer back in 5 minutes? Is it okay to wait for 2 hours?
Or do you need the answer back in 30 seconds? This goes back to kind of my AI factory research that I've been doing like GPU access is still relatively scarce. A lot of these problems are solvable on AMD epic and Intel Xeon and arm, and CPU is perfectly fine to solve some of these problems, but there has to be a reasoning layer that makes that determination for me. >> Absolutely. Absolutely. And then you if you add on HTA to high availability.
Because when we go talk to enterprises, they would say, "Well, I need my answers to come back between 30 to 60 seconds. " So, we talked about the problem. >> Yeah. You've modeled out your iceberg kind of framework for what enterprises need to be thinking about. Now, let's talk about the solution. Yeah. What are you folks helping doing it one internally and for your customers to solve this iceberg So, absolutely, we are customer zero for our products because of necessity, which is we have to enable as many people in the company to go as fast as possible and be able to build things and test things to their corner cases before it ever hits the customer.
The only way we can do that is use our platform and SDKs in ways that we ourselves haven't completely used up to now. Which is everybody is enabled to go use our platform, our SDKs to go build whatever they need to build. That's number one. And if you really think about it, the costs go up a little bit, but then that's in in terms of the returns that you're getting, it's nothing compared to the the actual outcome that uh we're enabling. That's number one.
" And this is what um drives the cost into your budget. " Like it thinks by itself and then it gives me a final output. We're like, "Yeah, but then you don't need to turn on thinking mode. " It's subtle things like that you need to get into the the mix. And most importantly, don't approach this as an add-on to your existing processes. Which is you have an existing workflow. Obviously, it was working, that's why your enterprise is working.
However, you need to think about what can you do with these tools that were impossible to do without these tools. Yeah, we were just talking to a town in Colorado that has 80 years of transaction of real estate transaction data on microfiche. There was a There's a human process for looking at up and referencing that. Instead of repeating the human process for doing that, they created a AI-specific process for doing it, and they're having success. Yes. And also, think about how you can push boundaries for yourself.
That's really what is big beginning to give the biggest unlock for us. And we're also starting to see that with multiple customers who you may think traditionally are very slow movers. Think of large-scale manufacturers. Think of major banks or securities agencies. Those are entities where the change management itself might take years. And they're actually moving much faster. So, as I think about this as a systematic change in how I operate, the great thing about AI, kind of the the ratchet here, is that I don't even have to learn the SDK.
I can point my ID, my code assistant to the SDK and the SDK documentation. In fact, we go one step further. We the build tools also come with cursor rules. Wait. You you just introduced something that I think forces me to like there's whole no-code platform designed for non-coders. >> Yes. Cursor's ID. It's as good as a non-code for a non no-code users. No, you're you're you're actually having these program managers are actually just using cursor? Yes.
They're like they're using cursor, that's what I mean. We didn't even build a tool that is completely no-code. It is the same tool that developers use. We don't have time to go build new internal tools. " Now, their usage of cursor is very different than a developer would, meaning they would never look at the code. >> Right. They don't have the big code Yeah, they don't care. The code editor The code editor could be window can be actually be closed.
But they have to chat window and then they're like looking at the chat and what is going on, but that is more than good enough for them to get to their outcome. So, as a IT security expert, as someone who's looking for guardrails, it becomes the rules themselves become the guardrails. >> That's right. That's right. And put a star there because yes, we have an AI coder internally that we actually ship with our product, but that coder can be cursor rules, it can be cloud code, it can be chemical code, any of those things you want.
We have AI code just because we can build our build tools and SDKs together, so it's all packaged together. But from an enterprise standpoint, what tools you particularly use to build anymore, as long as it hits the latest models, and as long as you have the right set of configurations to go with it, and you enable autonomy. That is important. Because that's another thing that changed in the last 8 weeks. Completely autonomous. " And it near autonomously builds something that I can then come back and look at.
And I think that's the big difference in the past 8 weeks from when we had the conversation. When we had the conversation, I was thinking more of your DSM in the terms of traditional uh professional services engagement. There's, you know, kind of your experts, there's your application developers or engineers, and then there's the customers, and the the engineers kind of just shim between the experts and the customer. Now, we've kind of uh uh uh of reduced that friction, and the experts and the customer are now sitting together and just Yes.
working through cursor. >> we did one more thing here, where we took our domain-specific model development workflow, which generates a massive knowledge graph as an input to building models, and turn that into what we call data perception. Which any enterprise can subscribe to to do on their own, which is you take you just point to wherever your data is sitting, whether it is sitting in SharePoint, Snowflake, Databricks, in your own like local object store, it doesn't matter. This will go in and it genetically process all of the data sets wherever they're sitting.
We as long as there is computer access to closest to the data, get you to a full-blown knowledge graph of your entire data set. And if you do it yourself, you just subscribe to the data perception app, if you may, and you get that. Or you can also get that as a service from us. The next thing that it translates to is once you have that, you can decide, do you need a domain-specific model to take advantage of all the core knowledge you have?
It'll also be part of the report that comes out of it. And then you can do model training with support from the platform, meaning it'll tell you whether you have sufficient data only to fine-tune, or do you have sufficient data to continuously pre-train, or it'll also say, "Look, you know what? You have enough data. " And these are the kinds of tasks you can actually do with it. So, I have to ask the question, cuz this is abstracting a lot of the kind of dirty work that that we have to do, and that I'm hiring my experts to do.
Where is the expertise moving to in the enterprise to do this stuff? So, the expertise really needs to be here to down, and also to make sure that you adapt to the changes that is coming at you. Because nobody can say, even the frontier model research labs can't say their models are the only things they can rely on. They have to be able to hit the other models to make sure how they compare. And it's not just the models are getting bigger because they're just getting bigger and more data set.
The architectures are fundamentally shifting. There's new things that are coming into the models that even with the research team, you need to have the ability to adapt. So, how does the platform, your platform specifically, help me solve or keep up with this problem? >> Perfect. So, first of all, your data is continuously This Do you call it perception because it's a continual process, so it keeps going on. But even the underlying models are fully supported by us, meaning every time there's a model update that comes in, anytime we update the model, which is by the way multiple times a week.
All of our subscribers gets all of those capabilities right away. So, I'm I'm focused again as a expert within the enterprise that's adopted your platform. I'm focused on my organization's goals, guardrails, security parameters, etc. And I'm not trying to keep up with what the latest version of uh uh of of Open AI's >> Sorry, the latest version of being saying you are what is the latest OCR model, none of that. Or even choosing between like even if I'm looking at uh the models from Meta, I'm not trying to determine which one of those models are the ones that I should be keeping up with.
>> no. I mean, in in fact, like Meta has gone so slow now that like say Llama 4 Maverick is getting deprecated in most of the model providers. Like because of all of the other models significantly leapfrogging. All right. Anything that I've that I've missed or we talked about a lot. >> of stuff. So, the other thing is like with domain-specific models, right? So, it's not that we used to be like say oh we are like 20% better, 30% better.
With the lot of the architecture changes and us using our platform to build our own models and change architectures, and I'll make a little bit of news here, which is we just finished a model that is our DSM for semiconductors. 6. Wow. Okay. Now, this is a 30 billion parameter model Mhm. competing against Claude knows maybe a couple of trillion parameter models, but in the domain of semiconductor design, in the domain of very large generation and validation, it's two times better than your frontier model.
Well, this opens up a whole another conversation around AI economics, etc. What happens when a a model is two x the the performance and a fraction of the size? >> Size and the cost and what they can actually I can run I can run that on a uh you know, we're we're pre-inve invited. I can run that on a 6000 Pro. Totally. Totally. All right, Arun. This has been a fascinating conversation. I know you'll be with customers having deeper conversations.
I'm looking forward to revisiting this conversation eight weeks from now. Maybe this whole thing will be obsolete. Who knows? We'll link the uh conversation that we just had mainly around reasoning eight weeks ago. It's super relevant because it tackles this part of the conversation about where our research is pointing to and what you need to understand and how uh uh Articulate is addressing the larger problem around not just the infrastructure, but the models themselves. com. com. Finally, at the 10 years.
Talk to you next CTO Advisor road trip light port conversation. Perfect. Thank you, Keith. That was fun. That was nice.