{"id":6863,"date":"2025-09-20T16:04:55","date_gmt":"2025-09-20T16:04:55","guid":{"rendered":"https:\/\/techtrendfeed.com\/?p=6863"},"modified":"2025-09-20T16:04:56","modified_gmt":"2025-09-20T16:04:56","slug":"faye-zhang-on-utilizing-ai-to-enhance-discovery-oreilly","status":"publish","type":"post","link":"https:\/\/techtrendfeed.com\/?p=6863","title":{"rendered":"Faye Zhang on Utilizing AI to Enhance Discovery \u2013 O\u2019Reilly"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"postContent-content\">\n<div class=\"podcast_player\">\n<div id=\"3202220113\" class=\"castos-player dark-mode \" tabindex=\"0\" data-episode=\"17467\" data-player_id=\"3202220113\">\n<div class=\"player\">\n<div class=\"player__main\">\n<div class=\"player__artwork player__artwork-17467\">\n<img decoding=\"async\" src=\"https:\/\/www.oreilly.com\/radar\/wp-content\/uploads\/sites\/3\/2024\/01\/Podcast_Cover_GenAI_in_the_Real_World-160x160.png\" alt=\"O'Reilly Media\" title=\"O'Reilly Media\"\/><\/div>\n<div class=\"player__body\">\n<div class=\"currently-playing\">\n<p>\nO&#8217;Reilly Media<\/p>\n<p>Generative AI within the Actual World: Faye Zhang on Utilizing AI to Enhance Discovery<\/p>\n<\/div>\n<div class=\"play-progress\">\n<div class=\"play-pause-controls\">\n<button title=\"Play\" aria-label=\"Play Episode\" aria-pressed=\"false\" class=\"play-btn\"><br \/>\n<span class=\"screen-reader-text\">Play Episode<\/span><br \/>\n<\/button><br \/>\n<button title=\"Pause\" aria-label=\"Pause Episode\" aria-pressed=\"false\" class=\"pause-btn hide\"><br \/>\n<span class=\"screen-reader-text\">Pause Episode<\/span><br \/>\n<\/button><br \/>\n<img decoding=\"async\" src=\"https:\/\/www.oreilly.com\/radar\/wp-content\/plugins\/seriously-simple-podcasting\/assets\/css\/images\/player\/images\/icon-loader.svg\" alt=\"Loading\" class=\"ssp-loader hide\"\/><\/div>\n<div>\n<audio preload=\"none\" class=\"clip clip-17467\"><source src=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3\"><\/source><\/audio><\/p>\n<div class=\"ssp-playback playback\">\n<p>\n<button class=\"player-btn player-btn__volume\" title=\"Mute\/Unmute\"><br \/>\n<span class=\"screen-reader-text\">Mute\/Unmute Episode<\/span><br \/>\n<\/button><br \/>\n<button data-skip=\"-10\" class=\"player-btn player-btn__rwd\" title=\"Rewind 10 seconds\"><br \/>\n<span class=\"screen-reader-text\">Rewind 10 Seconds<\/span><br \/>\n<\/button><br \/>\n<button data-speed=\"1\" class=\"player-btn player-btn__speed\" title=\"Playback Speed\" aria-label=\"Playback Speed\">1x<\/button><br \/>\n<button data-skip=\"30\" class=\"player-btn player-btn__fwd\" title=\"Fast Forward 30 seconds\"><br \/>\n<span class=\"screen-reader-text\">Quick Ahead 30 seconds<\/span><br \/>\n<\/button><\/p>\n<p>\n<time class=\"ssp-timer\">00:00<\/time><br \/>\n<span>\/<\/span><br \/>\n<time class=\"ssp-duration\" datetime=\"PT0H0M0S\">22m 12s<\/time><\/p>\n<\/div>\n<\/div>\n<\/div>\n<nav class=\"player-panels-nav\">\n<button class=\"subscribe-btn\" id=\"subscribe-btn-17467\" title=\"Subscribe\">Subscribe<\/button><br \/>\n<button class=\"share-btn\" id=\"share-btn-17467\" title=\"Share\">Share<\/button><\/nav>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<p>On this episode, Ben Lorica and AI engineer Faye Zhang speak about discoverability:  use AI to construct search and suggestion engines that truly discover what you need. Pay attention in to learn the way AI goes method past easy collaborative filtering\u2014pulling in many various sorts of knowledge and metadata, together with pictures and voice, to get a a lot better image of what any object is and whether or not or not it\u2019s one thing the person would need.<\/p>\n<p><strong>Concerning the<\/strong> <strong><em>Generative AI within the Actual World<\/em><\/strong> <strong>podcast:<\/strong> In 2023, ChatGPT put AI on everybody\u2019s agenda. In 2025, the problem shall be turning these agendas into actuality. In <em>Generative AI within the Actual World<\/em>, Ben Lorica interviews leaders who&#8217;re constructing with AI. Be taught from their expertise to assist put AI to work in your enterprise.<\/p>\n<p>Take a look at <a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/learning.oreilly.com\/playlists\/42123a72-1108-40f1-91c0-adbfb9f4983b\/?_gl=1*16z5k2y*_ga*MTE1NDE4NjYxMi4xNzI5NTkwODkx*_ga_092EL089CH*MTcyOTYxNDAyNC4zLjEuMTcyOTYxNDAyNi41OC4wLjA.\" target=\"_blank\" rel=\"noreferrer noopener\">different episodes<\/a> of this podcast on the O\u2019Reilly studying platform.<\/p>\n<h2 class=\"wp-block-heading\">Transcript<\/h2>\n<p><em>This transcript was created with the assistance of AI and has been calmly edited for readability.<\/em><\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=0\" target=\"_blank\" rel=\"noreferrer noopener\">0:00<\/a>: <strong>As we speak we&#8217;ve Faye Zhang of Pinterest, the place she\u2019s a employees AI engineer. And so with that, very welcome to the podcast.<\/strong><\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=14\" target=\"_blank\" rel=\"noreferrer noopener\">0:14<\/a>: Thanks, Ben. Big fan of the work. I\u2019ve been lucky to attend each the Ray and NLP Summits. I do know the place you function chairs. I additionally love the O\u2019Reilly AI podcast. The latest episode on A2A and the one with Raiza Martin on NotebookLM have been actually inspirational. So, nice to be right here.\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=33\" target=\"_blank\" rel=\"noreferrer noopener\">0:33<\/a>: <strong>All proper, so let\u2019s leap proper in. So one of many first issues I actually wished to speak to you about is that this work round <\/strong><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/arxiv.org\/abs\/2503.00619\" target=\"_blank\" rel=\"noreferrer noopener\"><strong>PinLanding<\/strong><\/a><strong>. And also you\u2019ve revealed papers, however I assume at a excessive stage, Faye, possibly describe for our listeners: What drawback is PinLanding attempting to deal with?<\/strong><\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=53\" target=\"_blank\" rel=\"noreferrer noopener\">0:53<\/a>: Yeah, that\u2019s an amazing query. I believe, briefly, attempting to unravel this trillion-dollar discovery disaster. We\u2019re dwelling by way of the best paradox of the digital financial system. Primarily, there\u2019s infinite stock however little or no discoverability. Image one instance: A bride-to-be asks ChatGPT, \u201cNow, discover me a marriage costume for an Italian summer time winery ceremony,\u201d and he or she will get nice common recommendation. However in the meantime, someplace in Nordstrom\u2019s a whole bunch of catalogs, there sits the proper terracotta Soul Committee costume, by no means to be discovered. And that\u2019s a $1,000 sale that can by no means occur. And in the event you multiply this by a billion searches throughout Google, SearchGPT, and Perplexity, we\u2019re speaking a few $6.5 trillion market, in response to Shopify\u2019s projections, the place each failed product discovery is cash left on the desk. In order that\u2019s what we\u2019re attempting to unravel\u2014primarily resolve the semantic group of all platforms versus person context or search.\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=125\" target=\"_blank\" rel=\"noreferrer noopener\">2:05<\/a>: <strong>So, earlier than PinLanding was developed, and in the event you look throughout the trade and different firms, what could be the default\u2014what could be the incumbent system? And what could be inadequate about this incumbent system?<\/strong><\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=142\" target=\"_blank\" rel=\"noreferrer noopener\">2:22<\/a>: There have been researchers throughout the previous decade engaged on this drawback; we\u2019re positively not the primary one. I believe primary is to grasp the catalog attribution. So, again within the day, there was multitask R-CNN technology, as we bear in mind, [that could] establish trend purchasing attributes. So you&#8217;ll cross in-system a picture. It might establish okay: This shirt is crimson and that materials could also be silk. After which, in recent times, due to the leverage of enormous scale VLM (imaginative and prescient language fashions), this drawback has been a lot simpler.\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=183\" target=\"_blank\" rel=\"noreferrer noopener\">3:03<\/a>: After which I believe the second route that folks are available is through the content material group itself. Again within the day, [there was] analysis on be a part of graph modeling on shared similarity of attributes. And loads of ecommerce shops additionally do, \u201cHey, if individuals like this, you may additionally like that,\u201d and that relationship graph will get captured of their group tree as effectively. We make the most of a imaginative and prescient giant language mannequin after which the muse mannequin CLIP by OpenAI to simply acknowledge what this content material or piece of clothes may very well be for. After which we join that between LLMs to find all prospects\u2014like situations, use case, worth level\u2014to attach two worlds collectively.\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=235\" target=\"_blank\" rel=\"noreferrer noopener\">3:55<\/a>: <strong>To me that means you could have some rigorous eval course of or perhaps a separate staff doing eval. Are you able to describe to us at a excessive stage what&#8217;s eval like for a system like this?\u00a0<\/strong><\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=251\" target=\"_blank\" rel=\"noreferrer noopener\">4:11<\/a>: Positively. I believe there are inside and exterior benchmarks. For the exterior ones, it\u2019s the Fashion200K, which is a public benchmark anybody can obtain from Hugging Face, on a normal of how correct your mannequin is on predicting trend objects. So we measure the efficiency utilizing the recall top-k metrics, which says whether or not the label seems among the many top-end prediction attribute precisely, and consequently, we had been in a position to see 99.7% recall for the highest ten.<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=287\" target=\"_blank\" rel=\"noreferrer noopener\">4:47<\/a>: <strong>The opposite subject I wished to speak to you about is suggestion programs. So clearly there\u2019s now speak about, \u201cHey, possibly we are able to transcend correlation and go in the direction of reasoning.\u201d Are you able to [tell] our viewers, who will not be steeped in state-of-the-art suggestion programs, how you&#8217;ll describe the state of recommenders nowadays?<\/strong><\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=323\" target=\"_blank\" rel=\"noreferrer noopener\">5:23<\/a>: For the previous decade, [we\u2019ve been] seeing super motion from foundational shifts on how RecSys primarily operates. Simply to name out a couple of huge themes I\u2019m seeing throughout the board: Primary, it\u2019s form of shifting from correlation to causation. Again then it was, hey, a person who likes X may also like Y. However now we truly perceive why contents are linked semantically. And our LLM AI fashions are in a position to motive in regards to the person preferences and what they really are.\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=358\" target=\"_blank\" rel=\"noreferrer noopener\">5:58<\/a>: The second huge theme might be the chilly begin drawback, the place firms leverage semantic IDs to unravel the brand new merchandise by encoding content material, understanding the content material straight. For instance, if it is a costume, then you definitely perceive its shade, type, theme, and so on.\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=377\" target=\"_blank\" rel=\"noreferrer noopener\">6:17<\/a>: And I consider different larger themes we\u2019re seeing; for instance, Netflix is merging from [an] remoted system right into a unified intelligence. Simply this previous 12 months, Netflix [updated] their multitask structure the place [they] shared representations, into one they referred to as the UniCoRn system to allow company-wide enchancment [and] optimizations.\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=404\" target=\"_blank\" rel=\"noreferrer noopener\">6:44<\/a>: And really lastly, I believe on the frontier facet\u2014that is truly what I realized on the AI Engineer Summit from YouTube. It\u2019s a DeepMind collaboration, the place YouTube is now utilizing a big suggestion mannequin, primarily educating Gemini to talk the language of YouTube: of, hey, a person watched this video, then what may [they] watch subsequent? So loads of very thrilling capabilities occurring throughout the board for positive.\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=435\" target=\"_blank\" rel=\"noreferrer noopener\">7:15<\/a>: <strong>Typically it sounds just like the themes from years previous nonetheless map over within the following sense, proper? So there\u2019s content material\u2014the distinction being now you could have these basis fashions that may perceive the content material that you&#8217;ve got extra granularly. It may well go deep into the movies and perceive, hey, this video is much like this video. After which the opposite supply of sign is conduct. So these are nonetheless the 2 foremost buckets?<\/strong><\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=473\" target=\"_blank\" rel=\"noreferrer noopener\">7:53<\/a>: Appropriate. Sure, I might say so.\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=475\" target=\"_blank\" rel=\"noreferrer noopener\">7:55<\/a>: <strong>And so the muse fashions enable you to on the content material facet however not essentially on the conduct facet?<\/strong><\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=483\" target=\"_blank\" rel=\"noreferrer noopener\">8:03<\/a>: I believe it is determined by the way you wish to see it. For instance, on the embedding facet, which is a form of illustration of a person entity, there have been transformations [since] again within the day with the BERT Transformer. Now it\u2019s obtained lengthy context encapsulation. And people are all with the assistance of LLMS. And so we are able to higher perceive customers, to not subsequent or the final clicks, however to \u201chey, [in the] subsequent 30 days, what may a person like?\u201d\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=511\" target=\"_blank\" rel=\"noreferrer noopener\">8:31<\/a>: <strong>I\u2019m unsure that is occurring, so appropriate me if I\u2019m unsuitable. The opposite factor that I might think about that the muse fashions can assist with is, I believe for a few of these programs\u2014like YouTube, for instance, or possibly Netflix is a greater instance\u2014thumbnails are essential, proper? The very fact now that you&#8217;ve got these fashions that may generate a number of variants of a thumbnail on the fly means you may run extra experiments to determine person preferences and person tastes, appropriate?\u00a0<\/strong><\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=545\" target=\"_blank\" rel=\"noreferrer noopener\">9:05<\/a>: Sure. I might say so. I used to be fortunate sufficient to be invited to one of many engineer community dinners, [and was] talking with the engineer who truly works on the thumbnails. Apparently it was all personalised, and the method you talked about enabled their speedy iteration of experiments, and had positively yielded very optimistic outcomes for them.\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=569\" target=\"_blank\" rel=\"noreferrer noopener\">9:29<\/a>: <strong>For the listeners who don\u2019t work on suggestion programs, what are some common classes from suggestion programs that typically map to different types of ML and AI functions?\u00a0<\/strong><\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=584\" target=\"_blank\" rel=\"noreferrer noopener\">9:44<\/a>: Yeah, that\u2019s an amazing query. A whole lot of the ideas nonetheless apply. For instance, the information distillation. I do know Certainly was attempting to sort out this.\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=596\" target=\"_blank\" rel=\"noreferrer noopener\">9:56<\/a>: <strong>Perhaps Faye, first outline what you imply by that, in case listeners don\u2019t know what that&#8217;s.<\/strong>\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=602\" target=\"_blank\" rel=\"noreferrer noopener\">10:02<\/a>: Sure. So information distillation is basically, from a mannequin sense, studying from a father or mother mannequin with bigger, larger parameters that has higher world information (and the identical with ML programs)\u2014to distill into smaller fashions that may function a lot quicker however nonetheless hopefully encapsulate the educational from the father or mother mannequin.\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=624\" target=\"_blank\" rel=\"noreferrer noopener\">10:24<\/a>: So I believe what Certainly again then confronted was the basic precision versus recall in manufacturing ML. Their binary classifier wants to actually filter out the batch job that you&#8217;d suggest to the candidates. However this course of is clearly very noisy, and sparse coaching information could cause latency and in addition constraints. So I believe again within the work they revealed, they couldn\u2019t actually get efficient separate r\u00e9sum\u00e9 content material from Mistral and possibly Llama 2. After which they had been comfortable to be taught [that] out-of-the-box GPT-4 achieved one thing like 90% precision and recall. However clearly GPT-4 is dearer and has near 30 seconds of inference time, which is way slower.<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=681\" target=\"_blank\" rel=\"noreferrer noopener\">11:21<\/a>: So I believe what they do is use the distillation idea to fine-tune GPT 3.5 on labeled information, after which distill it into a light-weight BERT-based mannequin utilizing the temperature scale softmax, and so they\u2019re in a position to obtain millisecond latency and a comparable recall-precision trade-off. So I believe that\u2019s one of many learnings we see throughout the trade that the normal ML methods nonetheless work within the age of AI. And I believe we\u2019re going to see much more within the manufacturing work as effectively.\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=717\" target=\"_blank\" rel=\"noreferrer noopener\">11:57<\/a>: <strong>By the way in which, one of many underappreciated issues within the suggestion system house is definitely UX in some methods, proper? As a result of mainly good UX for delivering the suggestions truly can transfer the needle. The way you truly current your suggestions may make a cloth distinction. <\/strong>\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=744\" target=\"_blank\" rel=\"noreferrer noopener\">12:24<\/a>: I believe that\u2019s very a lot true. Though I can\u2019t declare to be an knowledgeable on it as a result of I do know most suggestion programs cope with monetization, so it\u2019s difficult to place, \u201cHey, what my person clicks on, like have interaction, ship through social, versus what proportion of that\u2026<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=762\" target=\"_blank\" rel=\"noreferrer noopener\">12:42<\/a>: <strong>And it\u2019s additionally very platform particular. So you may think about TikTok as one single feed\u2014the advice is simply on the feed. However YouTube is, you understand, the stuff on the facet or no matter. After which Amazon is one thing else. Spotify and Apple [too]. Apple Podcast is one thing else. However in every case, I believe these of us on the surface underappreciate how a lot these firms put money into the precise interface.<\/strong><\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=798\" target=\"_blank\" rel=\"noreferrer noopener\">13:18<\/a>: Sure. And I believe there are a number of iterations occurring on any day, [so] you may see a special interface than your pals or household since you\u2019re truly being grouped into A\/B assessments. I believe that is very a lot true of [how] the engagement and efficiency of the UX have an effect on loads of the search\/rec system as effectively, past the information we simply talked about.\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=821\" target=\"_blank\" rel=\"noreferrer noopener\">13:41<\/a>: <strong>Which brings to thoughts one other subject that can be one thing I\u2019ve been serious about, over many, a few years, which is that this notion of experimentation. Most of the most profitable firms within the house even have invested in experimentation instruments and experimentation platforms, the place individuals can run experiments at scale. And people experiments may be carried out far more simply and may be monitored in a way more principled method in order that any form of issues they do are backed by information. So I believe that firms underappreciate the significance of investing in such a platform.\u00a0<\/strong><\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=868\" target=\"_blank\" rel=\"noreferrer noopener\">14:28<\/a>: I believe that\u2019s very a lot true. A whole lot of bigger firms truly construct their very own in-house A\/B testing experiment or testing frameworks. Meta does; Google has their very own and even inside completely different cohorts of merchandise, in the event you\u2019re monetization, social.\u00a0.\u00a0. They&#8217;ve their very own area of interest experimentation platform. So I believe that thesis may be very a lot true.\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=891\" target=\"_blank\" rel=\"noreferrer noopener\">14:51<\/a>: <strong>The final subject I wished to speak to you about is context engineering. I\u2019ve talked to quite a few individuals about this. So each six months, the context window for these giant language fashions expands. However clearly you may\u2019t simply stuff the context window full, as a result of one, it\u2019s inefficient. And two, truly, the LLM can nonetheless make errors as a result of it\u2019s not going to effectively course of that whole context window anyway. So speak to our listeners about this rising space referred to as context engineering. And the way is that enjoying out in your personal work?\u00a0<\/strong><\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=938\" target=\"_blank\" rel=\"noreferrer noopener\">15:38<\/a>: I believe it is a fascinating subject, the place you&#8217;ll hear individuals passionately say, \u201cRAG is lifeless.\u201d And it\u2019s actually, as you talked about, [that] our context window will get a lot, a lot larger. Like, for instance, again in April, Llama 4 had this staggering 10 million token context window. So the logic behind this argument is sort of easy. Like if the mannequin can certainly deal with tens of millions of tokens, why not simply dump all the pieces as an alternative of doing a retrieval?<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=968\" target=\"_blank\" rel=\"noreferrer noopener\">16:08<\/a>: I believe there are fairly a couple of elementary limitations in the direction of this. I do know of us from contextual AI are enthusiastic about this. I believe primary is scalability. A whole lot of occasions in manufacturing, not less than, your information base is measured in terabytes or petabytes. So not tokens. So one thing even bigger. And quantity two I believe could be accuracy.<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=993\" target=\"_blank\" rel=\"noreferrer noopener\">16:33<\/a>: The efficient context home windows are very completely different. Truthfully, what we see after which what&#8217;s marketed in product launches. We see efficiency degrade lengthy earlier than the mannequin reaches its \u201cofficial limits.\u201d After which I believe quantity three might be the effectivity and that form of aligns with, truthfully, our human conduct as effectively. Like do you learn a complete e book each time it&#8217;s worthwhile to reply one easy query? So I believe the context engineering [has] slowly advanced from a buzzword, a couple of years in the past, to now an engineering self-discipline.\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=1035\" target=\"_blank\" rel=\"noreferrer noopener\">17:15<\/a>: <strong>I\u2019m appreciative that the context home windows are rising. However at some stage, I additionally acknowledge that to some extent, it\u2019s additionally form of a feel-good transfer on the a part of the mannequin builders. So it makes us really feel good that we are able to put extra issues in there, however it might not truly assist us reply the query exactly. Really, a couple of years in the past, I wrote form of a tongue-and-cheek put up referred to as \u201c<\/strong><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/gradientflow.substack.com\/p\/structure-is-all-you-need\" target=\"_blank\" rel=\"noreferrer noopener\"><strong>Construction Is All You Want<\/strong><\/a><strong>.\u201d So mainly no matter construction you could have, it&#8217;s best to assist the mannequin, proper? If it\u2019s in a SQL database, then possibly you may expose the construction of the information. If it\u2019s a information graph, you leverage no matter construction you need to present the mannequin higher context. So this complete notion of simply stuffing the mannequin with as a lot data, for all the explanations you gave, is legitimate. But in addition, philosophically, it doesn\u2019t make any sense to try this anyway.<\/strong><\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=1110\" target=\"_blank\" rel=\"noreferrer noopener\">18:30<\/a>: <strong>What are the issues that you&#8217;re trying ahead to, Faye, when it comes to basis fashions? What sorts of developments within the basis mannequin house are you hoping for? And are there any developments that you simply assume are beneath the radar?\u00a0<\/strong><\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=1132\" target=\"_blank\" rel=\"noreferrer noopener\">18:52<\/a>: I believe, to higher make the most of the idea of \u201ccontextual engineering,\u201d that they\u2019re primarily two loops. There\u2019s primary inside the loop of what occurred. Sure. Inside the LLMs. After which there\u2019s the outer loop. Like, what are you able to do as an engineer to optimize a given context window, and so on., to get one of the best outcomes out of the product inside the context loop. There are a number of methods we are able to do: For instance, there\u2019s the vector plus Excel or regex extraction. There\u2019s the metadata fillers. After which for the outer loop\u2014it is a quite common observe\u2014individuals are utilizing LLMs as a reranker, generally throughout the encoder. So the thesis is, hey, why would you overburden an LLM with a 20,000 rating when there are issues you are able to do to scale back it to prime hundred or so? So all of this\u2014context meeting, deduplication, and diversification\u2014would assist our manufacturing [go] from a prototype to one thing [that\u2019s] extra actual time, dependable, and in a position to scale extra infinitely.\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=1207\" target=\"_blank\" rel=\"noreferrer noopener\">20:07<\/a>: <strong>One of many issues I want\u2014and I don\u2019t know, that is wishful pondering\u2014is possibly if the fashions is usually a little extra predictable, that will be good. By that, I imply, if I ask a query in two other ways, it\u2019ll mainly give me the identical reply. The inspiration mannequin builders can in some way enhance predictability and possibly present us with a little bit extra clarification for a way they arrive on the reply. I perceive they\u2019re giving us the tokens, and possibly a number of the, a number of the reasoning fashions are a little bit extra clear, however give us an concept of how these items work, as a result of it\u2019ll influence what sorts of functions we\u2019d be comfy deploying these items in. For instance, for brokers. If I\u2019m utilizing an agent to make use of a bunch of instruments, however I can\u2019t actually predict their conduct, that impacts the sorts of functions I\u2019d be comfy utilizing a mannequin for.\u00a0<\/strong><\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=1278\" target=\"_blank\" rel=\"noreferrer noopener\">21:18<\/a>: Yeah, positively. I very a lot resonate with this, particularly now most engineers have, you understand, AI empowered coding instruments like Cursor and Windsurf\u2014and as a person, I very a lot admire the prepare of thought you talked about: why an agent does sure issues. Why is it navigating between repositories? What are you  when you\u2019re doing this name? I believe these are very a lot appreciated. I do know there are different approaches\u2014take a look at Devin, that\u2019s the totally autonomous engineer peer. It simply takes issues, and also you don\u2019t know the place it goes. However I believe within the close to future there shall be a pleasant marriage between the 2. Properly, now since Windsurf is a part of Devin\u2019s father or mother firm.\u00a0<\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=1325\" target=\"_blank\" rel=\"noreferrer noopener\">22:05<\/a>: <strong>And with that, thanks, Faye.<\/strong><\/p>\n<p><a rel=\"nofollow\" target=\"_blank\" href=\"https:\/\/cdn.oreillystatic.com\/radar\/generative-ai-real-world-podcast\/GenAI_in_the_Real_World_with_Faye_Zhang_updated.mp3#t=1328\" target=\"_blank\" rel=\"noreferrer noopener\">22:08<\/a>: Superior. Thanks, Ben.<\/p>\n<\/div>\n\n","protected":false},"excerpt":{"rendered":"<p>O&#8217;Reilly Media Generative AI within the Actual World: Faye Zhang on Utilizing AI to Enhance Discovery Play Episode Pause Episode Mute\/Unmute Episode Rewind 10 Seconds 1x Quick Ahead 30 seconds 00:00 \/ 22m 12s Subscribe Share On this episode, Ben Lorica and AI engineer Faye Zhang speak about discoverability: use AI to construct search and [&hellip;]<\/p>\n","protected":false},"author":2,"featured_media":6865,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[55],"tags":[3726,5455,267,238,5456],"class_list":["post-6863","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-machine-learning","tag-discovery","tag-faye","tag-improve","tag-oreilly","tag-zhang"],"_links":{"self":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts\/6863","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=6863"}],"version-history":[{"count":1,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts\/6863\/revisions"}],"predecessor-version":[{"id":6864,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/posts\/6863\/revisions\/6864"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=\/wp\/v2\/media\/6865"}],"wp:attachment":[{"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=6863"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=6863"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/techtrendfeed.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=6863"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}<!-- This website is optimized by Airlift. Learn more: https://airlift.net. Template:. Learn more: https://airlift.net. Template: 69d9690a190636c2e0989534. Config Timestamp: 2026-04-10 21:18:02 UTC, Cached Timestamp: 2026-08-12 05:02:46 UTC -->