Algolia – Powering Discovery via AI

This podcast episode of the CTO Advisor podcast episode is sponsored by Algolia software. In this episode, Keith Townsend interviews Algolia's CTO, Sean Mullaney. The pair discuss a practical use for generative AI for e-commerce websites. The range of topics include AI Infrastructure, AI training, and AI Inferencing focused on the shopping experience. [...]

Transcript 3,280 words · about 22 min to read

Machine-generated from the episode audio and not hand-corrected, so names and technical terms may be imperfect. The audio is authoritative.

All right, you're listening to another episode of the CTO advisor podcast. I. I. I. to meet their business goals. On the other end of the line, I have the CTO of Algolia Software, Sean Mullaney. Sean, how's it going? It's going great, Keith. Thank you so much. Excited to be joining you today. You know, it's exciting to have you on. We kind of pre-gamed a little bit last week and I guess we got really into the weeds and we promise you we're going to try not to go into the weeds.

But I also promise you that we're going to have a podcast, a follow on editorial podcast where we talk about the weeds just to tease it. We're going to talk about how Sean and his team are in every cloud, how they're able to manage that from a CTO resource perspective and still maintain kind of their hybrid on premises colo cost model. Very, very good lessons learned. But until then, Sean, tell me, what does Algolia do? So Algolia is the world's leader in search and discovery.

So when you think about it, a lot of people, they think about search, they think about Google. Right. But actually, every single Web site on the Internet has a search box and every single Web site on the Internet kind of has to discover content, products, Web pages, these kind of things. And so we're kind of part of the foundational infrastructure of the Internet. We power search and discovery for 17000 Web sites across the across the Internet. And we actually serve almost two trillion Web requests searches a year.

So we're actually the second biggest search engine in the world behind Google. But people don't know our name because we sit behind the scenes and we help power the infrastructure for other companies. So as an infrastructure guy, I can appreciate this. You know, it's it's you know, you talk to a team about what do they need as far as infrastructure resources. And, you know, the eyes glaze over and say, I just want to run cold. The same thing happens with a lot of these customers.

I have a Web site. I have content on my Web site. I have a search box. I want my content consumers to come to my Web site, type in the five or six letter five or six word search term and find that term. They don't care about what it took for me to get that response to them in three seconds or less. What they care about is that they get that response in like two seconds or less. So talk to me about, you know, the challenge of delivering on that customer expectation when you're across such a diverse set of customers.

Yeah, absolutely. So, I mean, we serve customers. E-commerce is one of our big segments. Obviously, finding products for shoppers online is an important part of the Internet. We also serve like media companies like The Washington Post, for example. We do a lot of site search for companies like Stripe. They're great documentations, all powered by Algolia. So we do serve a lot of use cases. But I think, you know, one of the simplest ways to think about is let's take e-com as an example.

So when you go to an e-commerce site, one of the things we all love about e-commerce is there are so many different products you can choose from. There are a lot of like e-commerce stores where there are hundreds of thousands of products that you can search. And if you were to walk into a store, this would be like a football stadium sized store to show you all the products. Right. It would be entirely overwhelming. You wouldn't know where to start.

So obviously, the search bar is like one of the first places that people go to when they're trying to figure out and find the products they're looking for. We know that search has to be really, really fast as a starting point. We know that hundreds of milliseconds delay leads to customers abandoning the shopping experience. So Algolia has built its entire platform on scale and on speed. But the most important thing is in finding the most relevant items that people are searching for.

And one of the really like interesting things is when you think back 20 years ago, when the Internet first started, you know, the search experience hasn't changed very much. You know, you type in some keywords into a box. It then goes out and looks for Web pages or looks for products that has that have those words. And then it returns you, you know, 10 blue links from Google, still 10 blue links. The experience is very similar. I. and would love to talk a little bit about how Algolia is actually pushing forward the whole industry in terms of trying to understand customers.

And not just match keywords and give you links back. I need to know something about the customer for it to be better than a bunch of blue links. Yeah. I.? Yeah. I. is absolutely transforming the way that we search for things and the consumer experience of customers online. The biggest breakthrough is moving away from words is the thing that we match. So, you know, since the history of search engines, we've always been taking the exact word that you type into the box and going into the catalog, the index and trying to find all the records, all the Web pages, all the products that have that word in it.

And that produces OK results. But actually, human language isn't a great way to be understood. There are a lot of ambiguities. There are a lot of synonyms. When you're searching for things like problems to be solved or you want to find something and there aren't any results that match it. But there are things that are similar. Keywords are imprecise. And so the really exciting thing is we've known for about 10 years now that we can turn keywords into a mathematical idea called a vector.

And in essence, what a vector is, is it's taking the concepts and the meaning behind the words and it's mapping those words into a space. And we can search that space and we can come up with very similar concepts, very similar ideas, even if they don't contain the same words. So a very simple example is – oh, sorry. Go ahead. No, I was about to just give an example of one of the most frustrating searches I've done over the past couple of years.

I do this thing where I take a neon pen and I'm and I'm writing in the air. And it looks like I'm writing, you know, kind of I'm actually writing on a glass board, a white glass whiteboard. It's a light board, transparent and finding like a pre-manufactured glass light board is called. It's one of the most frustrating searches because the keywords really don't describe what it is that I'm looking for. So I try to be more descriptive, descriptive in the search like light board for writing for video production and writing in the air.

Just no results. And that context is one of the things that's missing. And I would give all kinds of money just to buy one of these things outright. I know there's pre-produced models of it. I just can't find it. Yeah, we've all had this experience where, you know, that you're asking a computer a question and you're basically asking a database to find something. So you're trying to cherry pick the words that you're using. And in that case, those words mean a lot of different things.

And it's probably bringing back a lot of, you know, things that have light or glass in them that just aren't the product you're looking for. Exactly. So in this case, we'd be able to match those keywords right into the vector that represents the concept behind what you're looking for. And one of the amazing things is, is we can actually handle over 50 different languages. So you could ask those questions in all these different languages and they would all map to the same vector.

So it's very language neutral. And we're able to then search for products and Web pages that contain those concepts that they may not actually contain the keywords that you've used. So these large language models, as they're called, have been trained on the entire, you know, kind of all the words and all of the language on the Internet. And they've picked up all of these ideas and they've seen the associations behind them and created like a spider's web of understanding. And so using these large language models, we translate all the products in an e-commerce store and then we translate the query.

And we're able to actually search the database for the same concepts, not just the keywords. So we're kind of moving from matching words to really understanding what you're looking for. So can it get to a point kind of meeting the expectations of not just the your customers, but their customers? Can it get to the point where I can go to a search engine and say, you know what, give me the components of a light board so I can build one myself? Yeah, exactly.

So if you were on a website, I'll give you a simpler example. If you went to a grocery store's website and you wanted to know how to make a certain recipe, you'd be able to ask for the ingredients in this recipe and name the recipe by name. And it could go out and it could figure out all the ingredients that are associated with that recipe based off of what it's learned, being trained on all of the recipes online. So what about concepts like hallucination, et cetera?

Are we going to see those types of problems creep up in our shopping carts? Yeah. So there are kind of like two different ways in which you can think about e-commerce shopping. One is a kind of a retrieval. So we're using these large language models to try to understand the products, understand what you're asking for. But we're only retrieving the products that are available from the merchant. It would be interesting, but probably not too helpful if we made up new products and generated new products that weren't available for sale, for example.

So you can think of what we're doing is we're using it in a search environment or retrieval environment. We are working on a product, though, which I think is going to be very exciting, which enables you to chat to a conversational commerce chat bot online. So when you think about it, when you have a very big catalog of items, hundreds of thousands of things to choose from, the first thing you want is some help. And, you know, when you go to a retail store on the high street, you know, there'll be an assistant there to help you.

You can ask questions, too. They could point you in the right direction for what you're looking for. And so we do think that these kind of conversational generative agents or chat bots are going to be very helpful in providing expert advice and guidance as you're shopping online. And I'm seeing these use cases where we're taking large learners, language models, putting them in very specific areas of expertise. And we see less hallucination because we're getting trained in that area. Like if I'm going to my local plumbing store, I'm not going.

Hopefully the plumber, the local plumbers are going to start hallucinating and talking about how to build a gazebo versus how to fix my faucet. I should be able to get that detailed knowledge. So when I'm asking them, you know, that little round hole that goes in a circular thing and they have contacts and it's not this huge, you know, unlimited knowledge base. It's very focused. So the the chances of hallucination is much, much reduced. Yeah. You know, we do a couple of things.

Firstly, we fine tune these models for our customers data. So if you have a plumbing supply website, for example, in your example, we would fine tune the model to understand the actual products that are available on your website. And the second thing we do is we very much restrict what the chat bots and the large language models are able to answer and the areas that they're able to provide assistance. So we restricted very much to the products that are on offer on the website.

Just as you don't want your plumber to be talking politics to you, you also don't want your e-commerce store to be giving you political opinions or advice or hallucinating. So the last question is, how is this delivered? I. and training models and the cost associated with all of that, I don't want any parts of it. How do you ease the adoption of this, you know, industry changing technology, frankly? Yeah, absolutely. I. models instantly without changing any code. So they can go into their Algolia dashboard and be able to enable it.

And for the folks who we're not working with, you know, they can come and sign up to Algolia and they can, particularly if they're like a small merchant, they can use Shopify. We've got great plugins to all the different e-commerce platforms. I. without them having to understand it or worry about optimizing it. That's really what we're there for. I. and it should be easy for them to be able to plug in and get the benefits without having to worry about hiring data scientists or trying to figure out the math behind it.

So, Sean, we have some dates on your calendar going into July of the time period where this is recording, where we're going to go in kind of deeper and talk about that back end, because you have to be kind of in all clouds to make this work. That's really expensive. And there is some cost sensitivity with the value that you're bringing. You have to be able to bring this solution to market in budget and have economies that kind of work out. And I want to dive into that part on the editorial side.

But before we go, is there anything else you want to share with our audience? Did we miss anything? Yeah, we can talk a little bit, I think, about the real breakthrough that we've worked on in Algolia. So maybe we talk a little bit about this hashing technology. Oh, you're not going to get you're not going to threaten me with a good time and get away with it. Let's talk about hashing technologies. So we've known for a while that vectors are a great way to get better relevance and to get better human understanding.

But the problem with vectors has always been that they're extremely large to store, which means they take up a lot of memory and they're really expensive to compare. They're mathematically very complicated. So until recently, they have been mainly used in small scale applications or in applications where you can spend a lot of money to run them. So we haven't been able to leverage them so far in real production environments where you have large scale, like big e-commerce stores, news sites, these kind of where you get a lot of traffic.

And so we've been on a mission for the last few years to try to figure out how we can scale these vectors up. And we've come up with this compression technology called hashes. So we take these extremely large vectors. There are thousands of dimensions in size and take up lots and lots of memory and we shrink them down to a very small size. So it's about one tenth of the size, but it still retains almost 100 percent of the accuracy. And we're able to then put these hashes in a traditional database environment where we can get the same levels of speed and scale and reliability that you expect from a database, but now empowered with these kind of AI vectors.

So we really are like the first company that's been able to launch this at a scale, at a speed, and most importantly, a cost that e-commerce companies find really attractive. So I think that's one of the many things I want to dive into when we go into our more technical conversation about how is that done from a practical perspective? Because if we compress vectors, if the vectors are then compressed, then there has to be some compute associated with using that hash and kind of let's call it hydrating the hash to actually apply it to the inference.

Or are we just, you know, you know, what tricks are you using to optimize the hashes so that just the relevant hash is used, the relevant parts of the vectors are used? There's a lot. There's a lot to talk about in that technical conversation, but I'm glad you brought up the concept of hashing, because one of the biggest problems with AI and adopting AI at any scale, especially for real time search, is that they're really, really, really heavy pieces of data. We're talking about algorithms with billions of vectors and potential input.

So I'm looking forward to having a technical conversation. Yeah, I mean, we leverage the same kind of large language models that power chat GPT. And if you've ever used chat GPT, it's amazing in terms of its power, but it's very slow, right? It takes a long time to generate those responses. And I know that companies that are providing them are spending enormous amounts of money behind the scenes to power these services. And so, you know, being able to address that level of scalability, speed and cost is really, I think, one of the critical things about bringing AI to everyone at web scale.

Well, I'm looking forward to that conversation. com. Links to Sean's work will be in the show notes. You want to ask me questions related to Sean, you can do that. At CTO Advisor on most social media platforms, my DMs are open. Sean, thanks a lot for joining us. Yeah, thank you so much. I'm looking forward to chatting further about it. All right. Thank you.