Your Data Is Getting Dirtier While We Talk — UnicornIQ’s Answer to AI’s Unsolvable Problem

Everyone says clean your data before you invest in AI. The problem is, your data is getting dirty faster than you can clean it. In this episode, Keith sits down with Joe Onisick, founder of UnicornIQ, to dig into why data hygiene is AI's most underestimated failure point — and why throwing more GPU power [...]

Transcript 3,478 words · about 23 min to read

Machine-generated from the episode audio and not hand-corrected, so names and technical terms may be imperfect. The audio is authoritative.

Alright, this is a treat. Joe, I think you're the first person, other than my family, that we've done a published podcast with in our home. Well, I appreciate you inviting me into your home. That is usually recommended against. You know what, you've had me at, what's the French name again? Name that for my mom, Avis Place. Avis Place, when there's a nice iron. You know, I'll put the picture as the title of the as the thumbnail. Pretty cool ranch.

You've had me there. You've had the CTO of Eyes of Flying Cloud there. This is the second time we've done a podcast, though, because we did a podcast there at the ranch. In which we ironically talked about data. We did. We did. And we're talking about data today. You're sponsoring this podcast. We're talking about your startup Unicorn IQ or UIQ for short. And you're solving where I think is a very unsolvable problem. Great garbage in garbage out.

Everyone tells us that you need to clean up your data before you put it into a. The problem is, is that once you clean it, it gets dirty again. Absolutely. It's getting dirty in real time. While we're talking, your data is getting dirtier. We're building technical debt. 80 percent of, you know, we've seen the numbers. I've heard CTOs, Chief Data Officers talk about this number time and time again. 80 percent of their data scientists spend. I'm sorry.

100 percent of their data scientists spend 80 percent of their time cleaning data instead of analyzing data. Now, everyone in our organization uses a chat bot and their pseudo data scientists. Absolutely. Absolutely. Because if you just slap a very expensive piece of computing over dirty data, eventually, if you throw enough money at the problem, it'll fix itself. Right. Yeah. We're infrastructure people. So you know what? The application is too slow. You throw more CPU and memory at it.

The application goes fast because CPU and memory is cheaper than developer. However, this is not the case. GPUs are quite expensive. CPUs are expensive. This whole inference pipeline is expensive. And even if we could throw money at the problem, we may not be able to keep up with the problem because the data is growing faster than we can throw CPU and GPUs at it. Agreed. And it's hard to think around the problem because all the marketing is saying just wait for the rag to fix it.

Wait for AI and inference to fix it. But I think people forget that incentive drives behavior. Who's selling rag? The people that are charging you by the drip for it. Of course, they want you to use that to fix it. If I'm selling you fuel, I want you buying Ferraris, not Corollas. This is true. So. On top of that, they abstract away the problem so much and their models are improving. I would say exponentially that there's this false sense of the problem is solved, but we see the cracks all the time.

Right. Right. Yeah. So I think when you look at it, an engineer would say that we've moved into a nondeterministic behavior. But what that really means is that I can't trace and troubleshoot AI's answer back to exactly why it gave me that answer. It's doing some level of magic that isn't completely traceable in that inference stack. Right. So this creates a problem. Think about an employee that works for you. Right. AIs are official intelligence. We have human intelligence.

So you're working with human intelligence and you ask me on your staff, hey, Joe, what should we do next with our marketing launch? And I say, um, we should probably run naked through the streets of Manhattan. And you go, well, why? And I go, you know, I don't know. It just sounded like a good idea at the time. That's AI. Right. I'll give you an example of how AI inference works. If I ask it to build me a go to market strategy for unicorn IQ, it'll build me one.

And if my next question is, show me your work, tell me how you came to that conclusion. Any AI prompt I type that into will immediately start lying to me. It'll start trying to provide human logic for how it made its decision when the real answer is I predict the next word. And then I predicted the next word. And then I did that a billion more times. And that is the. That's both the brilliance and the frustration with AI, with generative AI.

It looks like intelligence, but it's not intelligence. And what aspirates the problem or exasperates the problem is the concept that once I put garbage in that predictive nature. So if it's inferencing off of bad data or inconsistent data, we like to use the example of a Ford truck that, you know, 1935 Ford truck came in black, black only. That was an inspection. That is truth. Fast forward to 1975 or 1985. That truth differs. And when you give this to Iraq system, what's going to be the result?

Absolutely. So the result is not just the hallucinations we're getting. So there are all the scary numbers out there. We've talked about them. Eighty five percent of AI products fail. Ninety five percent of projects fail. Nobody's going to get a real number because there's nobody that goes out, reports their failures. Right. You know, that's not a thing we normally do. But the fact of the matter is, even when AI projects aren't failing, they're generally not getting us the result cost benefit that we were expecting out of them.

Maybe they're stalled versus failing. What's happening is a few things. First, we're making a smart assumption that went wrong when the context change. The smart assumption is I use chat GPT or whatever your AI flavor of choice. They're all comparable. It doesn't matter. I use that for public knowledge and it's awesome. Right. Like I can write like Twain. I can think like Plato and I can math like Einstein with AI. That's awesome. So I should be able to apply that same thing to my corporate data and get the same result.

Different data set. That's where the thing breaks down. First of all, the AI itself is trained on the entire corpus of human knowledge. It's been digitized. That's the largest data set ever available in our species history. Your data, in comparison, is one one billionth of that size. So you're already searching for a needle in a haystack every time you do inference. And then that data is dirty. I jokingly and lovingly say that if I went and ask, it's Benioff in charge of Salesforce.

Right. If I if I caught Benioff in a bar and asked him, is your CRM up to date and accurate right now? That the CEO of Salesforce would tell me, no, no, no, no. If he was being honest. Right. Our data is just continuing to get bad. So that's always churning and rag. We're always working harder for no reason, even if it's giving us decent results. We're paying way too much for those results. So let's talk about how you solve that problem.

Give me the elevator pitch. How are you taking, which is a problem that we've had since the first time we've said save data to a bit to disk has been the problem with knowledge management. How are you helping solve that problem? Mainly begrudgingly, Keith, this was not the problem I originally set out to solve. We were trying to build a higher level sales intelligence system. And what we realize is that to get to that level, I've got to bring in my customers data.

So we started doing that. We started testing. Everything fell apart every single time because we couldn't get accuracy close enough to be a usable system. People will continue to adopt because the underlying data was dirty. So we started looking at that problem and thinking about it. Now, when people talk about data hygiene, I believe that they overgeneralize it. They think of it like a light switch, like I need my data is dirty now and I'm not going to disagree with that too much.

But I don't have the money to switch the lights on and make it all clean. Right. Would you mind your time? Right. I agree. But. While you can't just pay for switching a light, money or time. Do you really want to knowingly invest hundreds of millions of dollars in an AI stack that's working on a data set you know is bad and getting worse? That just seems like a stop if I can't. So this is we've seen this play out time and time again with cloud, big data, you name it.

The once we figure out that once we think we can out engineer our bad data, our projects typically fail. So you're helping to not clean up the data, but help me discover truth. That's right. So what what we're doing is, first, we don't want to touch your data. Your source of data is the is the closest thing to fact that exists, even when it's wrong or right. I don't want to be the one making up new wrong out of it.

So my job is to sit out of the line of your data. You can still have whatever rag inference, BIA, any of the systems you use. I don't get involved with them. I create a data plug in that can ingest today from unstructured data. Like like you did files we use on a daily basis as knowledge, order PDFs, documents, the stuff that makes up 80 percent of your rag layer. I parse it, distill the information out of it, and then my system goes through a series of mathematical calculations to find out how confident we are in that data.

Now, some of its rocket surgery, some of its basic. Right. Let's say you have nine out of 10 to disagree. Right. That commercial we've all seen. Well, great. I can probably trust nine dentists more than I can trust one. But it doesn't mean one is wrong and it doesn't mean nine is right. It means I'm higher confidence. Right. So we're only ever assigning confidence based on what we can see. We're not saying this is fact or false.

We're saying same thing is if I ask Keith a question, he gives me an answer. I say, Keith, how confident are you? You might say 80 percent. Cool. Great. Now I know how much I can trust Keith's answer. That's what we're handing to the inference layer. I. you're prompt. Right. Like the thing you're trying to do, ask the question. It first is prompted to call the source of truth. The source of truth provides back a list of facts that exist in your data with a link to the original source that you can find that exact word, that exact sentence, that exact fact and trace provenance back to your data that I haven't modified.

And then serves that to the inference layer, which now gets to say, OK, I actually have facts on this subject. And these three are the highest confidence. So let me build my answer around that. The interesting thing is that by moving this left of the stack, I can do it deterministically. I can meaning I can give you results we can verify. And I end up optimizing the stack to the right side of me, meaning your prompts respond faster. The quality of response is measurably better and you're burning less CPU cycles and tokens to do it, which means reduction in cost.

Oh, by the way, if you're running that privately, that's also a reduction in power, cooling and fresh water for the cooling. So I'm hearing I'm creating this rich data set, this metadata around my data, which isn't in itself. A new concept, what is new is applying this to my data pipeline. That's a lot of work that I've put into creating this metadata registry to help me increase the confidence of my output. Who owns that now created metadata? Because now this is actually quite valuable.

I cannot just use it in my pipeline is I assume I can use it in all other types of business analytics. What's gaining me from either taking that system and applying it to other systems within my environment or even taking my ball and going somewhere else? Nothing, we're actually designed to be a completely open module. So if you if you think of me as a data hygiene black box and I use black boxes analogy, I'm happy to tell everybody what happens inside the box.

But on one side, I have data plugins for any supported data set. Right. So right now, G drive is one of the ones I support. So I can reach out to G drive and pull the files and data out of there. I can parse them and distill the information. And then I create a vector database, which is just a source of truth. It's not really even metadata about your data. It's the facts in text with a vector index in a gravity weighted database.

I expose all of that via API and MPC server. So I'm designed to integrate with whatever stack you have and sit in the back. So I'll give you an example. I. sales chatbot that just isn't getting traction because it's not giving right enough answers often enough. You've already invested a ton of money into a rag and inference stack sitting there. You don't want to rip and replace that. Not for me, not for anybody. Don't just add a prompt to that existing agent and tell it to check for truth first.

You'll instantly get better results across that entire stack because you're grounding in confidence, even in the noise. And that's how virtual CTO advisor works. The is grounded in my data set. My data set is relatively small, only a few thousand instances. But I even have problems in that because truth has changed over the 10 years that I've been writing content publicly. I. no longer applies in 2026. But that data is there. So you're helping me solve that problem.

Even in my small data set, helping me establish truth or confidence and then using the larger models. I will challenge you here, though. What happens when the data or the answer doesn't exist? Because that happens in my data, my data. I say, well, the answer doesn't exist. If someone goes goes to virtual CTO advisor and asks you about a topic that is not trained in or not aware of, then it will hallucinate. Absolutely. Absolutely. So let me let me let me start with an analogy.

As we talked about, I own a cattle ranch. And when I bought that cattle ranch, I started learning things I've never done before, like raising cattle, cash trading bulls and wrestling cattle out of hay feeders and all sorts of fun stuff. I also started investing in heavy equipment. So I a year ago bought a bobcat piece of heavy equipment. I want one. Yes. It comes with a manual. And that's great. The manual tells me how to use it.

But I can use that machine any way I want. And I often use it in ways that are very dangerous and not prescribed by the by the manufacturer. Our system is designed to use it however you want. We have some prescriptions for how you use it as to what you're trying to get out of it. But it's designed to be a very flexible, modular system, so it can plug and play into a lot of different use cases. I'm taking the long way around to the answer is it's going to depend on what you want out of your system as to how it handles a question it doesn't know the answer to.

So I'll use sales as two examples. If I'm a pharmaceutical rep talking to doctors and I say the wrong thing, consequences, major consequences. So in that case, if they're asking a chatbot a question, the chatbot doesn't have truth to they should get no answer and they should give the customer no answer for any amount of time until they have the right answer. Now, if I'm selling Cisco switches, I might not have the same worries like it's not quite the same. Right.

You know, if I accidentally say it's 52 ports, which is not a real thing, by the way, but instead of 48. Long story short, that's a different type of brawl, right? If I'm launching nuclear missiles, that's a whole nother level of fact I need. So our system is designed to be integrated into what it's trying to do. So for a general sales team, when they ask a question that's not supported in truth. They can click a button that says ask an SME that'll trigger a human in the loop session chat session generally.

When that SME answers, we'll actually record that as a digital fact, right? We'll create a new artifact out of that accepted answer, send it back into the pipeline. So next time the question gets asked, it doesn't have to be answered by a human. But that said, that human could still click a button and say just give me inference anyway. Right. I. to theorize about where the industry's going, wrong answers aren't a problem. You're brainstorming, right? I. to get you your brain thinking.

So in that case, great inference, even if you don't have facts, it depends on the situation, what's going on. So we configure it to work that way. But the magic of what we did is not ignore human in the loop because everyone thinks it's a dirty word. We built human in the loop and is a feature. So I'm taking friction away from a specific job task while pulling knowledge out of their head and building it into digital truth. Last question is about commercials.

If I don't ask you about how I pay for this thing, my audience will kill me. So what is this? Is this SaaS? Is this on prem software? Am I charged by the number of connections? Like what's the commercials about? So, Keith, this one's hard to explain and I don't know how far I'm allowed to say. T. for 30 years. When I started, we used to solve problems by building systems that solve the problem. I've spent the last 15 years helping companies take garbage to market.

It's lock-in for the sake of lock-in. It's cost for the sake of cost. It's SaaS, not because SaaS is the right delivery model, but because SaaS gets me ARR and that's what my VC wants to see, right? We're not making decisions based on anything but money in this industry right now. I decided to do the opposite. This is a modular, open, integrable system. I deploy it in your stack. You own it. You take care of it. You don't trust me with your data because you shouldn't trust anybody with your data.

And I'm not asking you to make me an exception. And I'm going to charge you flat-banded pricing based on the ingest cost. No surprises, no nothing, extremely fair pricing down to the CTO advisor and up the ExxonMobil or whoever. Well, I really appreciate you spending the time coming and talking to me about Unicorn IQ. If this just whet your whistle about the topic and you want to dive deeper, we actually did a light board. We did. And the massive CTO advisor studio basement.

And we talk through just the overall architecture and the philosophy around the product. The link will be below to watch that video or in the show notes if you're watching this or listening to this via your favorite podcast. If you enjoyed the conversation, click subscribe. Learn more about Unicorn IQ over on the website, which is? ai. That's a proper AI website. You want to learn more about the CTO advisor? com. com still works. com. Until then, see you on socials.

Talk to you next CTO advisor podcast.