From Data to Decisions: Futurum’s Intelligence Platform and GPT-4.0

In this episode of the CTO Advisor Podcast, Keith Townsend, a Global Advisor at the Futurum Group, and Erik Bethke, the CTO of Futurum, discuss the latest release from OpenAI, GPT-4.0. They begin by sharing their experiences at Futurum, where both have been working for seven months, and highlighting the development of a new intelligence [...]

Transcript 3,500 words · about 23 min to read

Machine-generated from the episode audio and not hand-corrected, so names and technical terms may be imperfect. The audio is authoritative.

Yep. All right, you're listening to or watching, depending on how I decide to publish this, another episode of the CTO Advisor podcast. This has been a long time coming. I've been with Futurum for seven months now. Eric, how long have you been at Futurum? Also about seven months. Yeah, so we've both been at Futurum for seven months. The Eric I'm referring to is Eric Bethke, our CTO. So for the first time in a long time, at least the past seven years, I'm working at another company that has a CTO.

Previously was just the CTO Advisor. But to be fair, Eric, you've been fairly busy. 0, O is an Oscar. That was funny. Before we get into that, what have you been working on at the Futurum group? Yeah, so I'm most closely working with the intelligence group at Futurum, building out a whole new platform for our intelligence dashboards and insights. It's really, really pretty cool. We're redoing the product from the ground up, all our own code, all in-house, and really making love to the product and making sure that we're delivering those insights to our customers exactly as they need them and want them.

It's going well. We just started releasing it. And what's pretty exciting is, honestly, they don't spend that much time with the friction on the app at all. They just are actually really focused on the data and the insights. And they're really thirsty and hungry and tearing through that. And that's really validating. And it's one of the things that I found interesting. I've gotten a few previews. I actually have access now. So I'm going to go in and see if I can find various insights around AI.

It's one of those things that, as I've used other analyst platforms, you had to download the data and then get insights yourself. And our intelligence platform is exactly that. It provides intelligence. It summarizes insights from the data that we collect and provide. Yeah. And we're just getting started. We're just getting started. We're going back through the back catalog of all of our practices, really grooming the data, refreshing the data, going through and normalizing it, and making sure that we have the most easy-to-grok visualizations for each subcategory and making sure that's all clean.

And then we have a very, very robust roadmap of features that we're going to be rolling out. But that's probably not what this podcast is about. No, it's not what this podcast is about, and we'll get onto that in the future. 0. And I'm over here at Smartsheet Engage in London. So I haven't gotten a chance to really dive into the announcements and give an overview and be able to get a smart overview. 0 and GPT-4? Yeah. If I had to distill it down to one single word, it's engineering.

And it's confusing the Twitter X space so much. And let me try to start to unpack that. Before this event was announced, Sam Altman started teasing us that something's going to happen. And people are like, oh, are we going to get QSTAR? Are we going to get that fabled agent that's going to be able to decompose long-running tasks? 5, right? And then last week on Hugging Face, there was a cute new model, GPT-2, rolling around. And Sam Altman subtweeted and said, oh, I'm pretty fond of GPT-2.

So there was all kinds of speculation, anticipation around this event. O. So the timing couldn't have been more designed on their announcement. But going back to that word engineering, they did not release QSTAR. 5 or chat GPT-5. And so that caused a lot of people to be disappointed, especially those that are really deep on the AI side and are looking for truly new frontiers on the model. But I am blown away. I am deeply impressed with OpenAI, and I would like to walk through it.

And I think it's extremely impactful, not just to the folks that are deep into AI or not just big tech, but I think they're very, very challenging to everyone, indies across the board. So first of all, let me just drop some specifics. 0, it's O for Omni. This is not a text model that dynamically pops open an image generator when you want an image, or it's not a text model that is wired up to the Whisper API or 11 labs for voice back out.

It is Omni because it's natively, natively a text image audio model all combined. So let me think about this from the context of traditional collaboration. Yeah. If I'm sitting in a room with you and we're thinking about our intelligence platform, like Eric, I really want a workflow that helps a customer from data to decisions. And you would naturally get up and draw that. You wouldn't say, let me email you a description, et cetera. You would just get up and draw it.

So is it in that vein that it's that Omni view that we're thinking about of, how would a intelligence interact in real life? That's right. That's right. They had a really good live demo that described just that. One of the demo folks drew out a simple linear equation on a piece of paper, three X plus one equals four, and asked GPTO to help tutor me, help me understand this, but don't give away the answer. So they used a Sharpie and a pen and they drew out that equation and just used their mobile phone and GPTO was able to see that video stream.

And it sliced images off of that. And with those images to read that equation, three X plus one equals four and verbally assist this person to go through and solve it. Now we've had audio, we've had whisper. What's different is a 300 millisecond latency, a third of a second. That is a game changer. The best in class we had before this was also OpenAI's Whisper API at three seconds. Everybody's used voice assistance and they're, they have been dumb and frustrating because you start talking and they get confused.

They listen to a little pause in your sentence and they just start talking over you and you get irritated and you have to stop, stop, stop. You have to hit the button to say, stop. No. 0, it actually has native emotional sentiment analysis, which you would think, well, what does that really matter in an enterprise context or something? I'm running a business meeting. I'm right. I'm working on the whiteboard. I'm trying to solve some problem. Why do I, what do I care if it has this emotional response and what do I care if it's interruptible?

Because that's just key to how we work as humans, how we collaborate with each other. You know, you have the ability to come over the top and interrupt me right now and I would catch it, you know, and it'd be graceful and I'd understand and back and forth. So it has now taken this native multi-modal capability and has delivered a really, really important human machine interface use case. So one of the things that, you know, organizations that we talk to get frustrated with when it comes to chat agents or models, let's not call it chat agents.

Specifically, they get frustrated at models because the models are being presented in chat format. Chat is just the application that the model is using. So ChatGP, of course, in open AI, most popular use case is to be presented to us, the user in the interface via chat. How does this change? How does this Omni model change other types of applications we might want to power? Like how does this engineering change take effect? All right. That's a really good question.

And that's part of what's so impressive about what they released. So real time, absolutely low latency Omni voice. I've talked about that a lot. They also delivered it being four times faster than GPT-4, four times faster. You know, if most tech companies get pretty excited, you know, if they release something 50 percent faster or two times faster, but four times faster, that's a big deal. They also, at the API level, drop their price again in half. You asked about building something.

0 model. Why? Because they're really good at writing their APIs. Their documentation is clear. It's stable. It's just performant. You know, one hour after their announcement, they've released it during the announcement. They released the new web app. They've updated their API and they started rolling out the desktop application, which I haven't chatted about yet. But the ability to do an announcement and roll it out across the chat application, plus the playground, plus the API and not suffer a hitch on scalability performance.

And at the API level, that means we can wire up whatever we want to wire up, and now it's four times faster, twice as cheap and low latency. It's this is a stuff. This is stuff that sets them apart from basically everybody else. No one else has that breadth of engineering chops on the delivery of the service layers. Yeah. So now I'm starting to think way beyond chat as I think about batch processing things that I may not have fed to ChatGP-4 because it was either too expensive or too slow or both.

Now that it's four times faster and it's half the cost, I can start using it in applications that I may not have used it before. And since it's omnichannel, you know, I'm thinking about video or real time video and being able to say, hey, you know, I come from regulated environments and now I can start listening in on customer calls or environments where I'm making a trade or I'm having a discussion that falls within the regulatory bounds and, you know, automatically starts recording or automatically starts watching the trades that are putting more attention to the trades that I'm making to ensure that I'm not going foul or regulatory challenges, which are things that while I might have been able to build with GTP-4 may not have been cost effective or may have slowed up the business process in a way that was unacceptable.

That's right. That's right. All sorts of subtle use cases come out of it, like recording this podcast here. You could do things with the sentiment analysis, like say, hey, review where we kind of like bogged down and we lost a little bit of excitement, where there is a little dead space, you know, or give me a 60 second snippet of just the highlights where we seem to really zing. You could do that on customer calls, podcasts, you know, perhaps if you're a company with a big body of video commentary, there's all sorts of use cases that come out of that.

So I'm interested to see where you see the overall market going. You're not just, you know, we kind of we at Futurum, you're super smart and we kind of stick you in the room and say, hey, build our stuff. But you're also obviously well versed about the overall industry because you're a user, you're a builder, you're building this stuff and you need the best tools to build the best product. Where do you see the mold that open AI is continually to build and where challengers will really have a problem overcoming?

Yeah, that's an exciting question. I've been thinking about this a lot since yesterday. Overall, I'm going to sound like a fanboy of open AI and I'm very impressed with them and I really do like their products and services. But I'm also simultaneously calling out kind of a warning shot to everybody. People really need to pay attention. 0 for free. I noticed that this morning. So you get a chat GPT-4 class model for free and it's wickedly fast. So, for example, the people who are challenged GrokQ, I really love those guys, really impressed with their inference dedicated chips and the speed of Grok serving Llama3 was really impressive.

You could get 500, 800 tokens per second. So you could get a wickedly fast text model. But now open AI says, well, how would you like something also wickedly fast, also free, but also omni? You know, that's got to be tough. You know, I just saw a researcher from Llama3 posting late last night on Twitter talking about they're very confident that they're about two months behind on delivering on the same omni capabilities as open AI. So the pace of change is so fast right now.

So one of the most dominant things is how does an enterprise or a company even keep up with the technical pace of change? Never mind the consumer, never mind society's adaptation to AI. How do you practically stay up to date on the pace of change? That's one thing. That's one of their moats. It's just raw executional speed. Another moat is by giving it away for free. This is subtle. They didn't talk about it. You know, they get to use the marketing language, the aura of bringing, you know, next generation AI intelligence to the masses for free out of the generosity of their heart.

And I know from talking to them, you know, they genuinely believe that. They're generous, good people. But there's a business purpose here. You get to collect all these voice prompts from the free users, and you get to fine tune and improve all of your human feedback on the voice layers. That is a moat. Just like Midjourney, when they do image generation, they have the best image generation in the world, and they've got that four little response on Discord, one, two, three, or four, which image do you like?

The most precious resource going forward is going to be well-labeled human data. With well-labeled human data, you can keep improving your AI products. So whoever can build up a flywheel where people are enthusiastically throwing their well-labeled human data at that solution is going to win. You know, because the fundamental AI piece, this is crude, and I'm a little bit dramatic here, but the actual model making is is kind of a commodity. You know, we used to be excited with 100,000 token windows.

Now we've got a million token windows. We used to be excited when they could put together a poem, or a song, or something like that. But they all can do that. They all can, you know, generate really, really impressive text. All of them can help you with coding. You know, the other thing that they dropped, I mentioned it lightly, is a desktop app. It's coming out for Mac first. It'll be a PC second. 0 lives with you in your desktop.

And you can share it with your screen and use your voice. You're like, hey, I'm stuck here on this coding problem. Help me refactor this file. Or, you know, I need to solve for some flights. Or I'm at the US Trademark Office's website. I don't quite understand the forms I'm supposed to fill out here for a trademark. Anything you see, it can digest the images, and it can give you back voice and solutions at that near real-time 300 millisecond latency.

It's extremely disruptive to so many businesses. Dramatically, Duolingo stock dropped 4% during the announcement yesterday, because it also can do real-time translation and learning. It's very disruptive. And here's kind of like the chilling part. 5 or 5. They kept all their dry powder in their back pocket. They come out four times faster, 50% cheaper in the API, free for the four users, a desktop application, rolled it out simultaneously on the API level, the web app, the desktop. They are just clearing the decks and raising the bar that high and saying, hey, if you want to be an AI company, this is what you need to deliver casually on a spring update.

Yeah, I'm getting really excited. I'm going to start refreshing my app page, too, because the idea that I can replace... Actually, I tried Gemini on my Google account, and my main complaint was that it wasn't as good as ChatGTP 4. And I don't mean there's a thing I have to find the language for it, but this friction that exists at being able to answer the emails or create the documents or create the content that I'm looking to create, it just seemed Gemini was not as intuitive as I would hope.

You have all the history of all... You have all of my data. You have my style. You should be able to answer an email in my voice. For really important communications, I still find myself copying and pasting the portions of an email or going to ChatGTP to help me communicate the difficult thing that I want to communicate versus the native tool. So the idea of having that across my entire desktop, that's pretty exciting, and that's as far ahead of the curve as you can get.

Yeah, yeah. You talk about moats and try to think about different competitors. I mean, the following is all just speculation on my part, but I'm really impressed with Meta. I'm really impressed with the strong effort led by Yann LeCun on open sourcing the Llama models. Yann believes strongly that the text-based models will not lead to super AGI, and that's why he's relaxed and that's why he's comfortable advocating the open sourcing of the Llama models. And then frankly, Facebook doesn't make money off of just inference off their AI model.

So I'm bullish that Meta will be able to continue to prosecute their mission, and plus they have a very, very strong founder that has authoritarian control of his company, and he can just do stuff like buy more GPUs than the rest of the planet combined. So Meta's got something strong going on. Google's got Demis Hababi on DeepMind. Fantastic guy, also from the game industry. I'm a huge fan. And what they've just, you know, they've just released the AlphaFold 3 is so good for humanity and medicine and research.

But I'd like to see that group let off the chains. I'd love to see Google release a bit faster, move a little bit quicker. I feel like that org is so careful that it is leaving a lot of opportunity on the table. Well, it is Google moving fast in AI when you have the size of Google and you have to protect search. I kind of wish the Alphabet did it similar to Google Cloud, make it its own thing and, you know, have sanity checks, but they do really need to take that dedicated leader, make them a CEO of that group.

This is just me giving advice to Sunar. I'm sure he's listening to the podcast and he's taking, you know, notes, fiercely taking notes on what two geeks think he should do with his trillion dollar business. But we digress. Eric, I really appreciate you jumping on. This has been a long time in coming. We've, you know, said, oh, yeah, we need to get on the horn. 0 that got us to actually get on a video call, have a conversation, share our thoughts at large with the Futurum Group and CTO Advisor audience.

com. com. Yep. com. I'm on X as CTO Advisor Eric. I'm on X, too. I am not the biggest Twitter guy. I only got like, you know, 800 people on me, but I'm at Bethke Eric. All right, Eric, thank you for joining. Talk to you next CTO Advisor podcast.