
Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
AI Executive Summary
Google Research scientists Nitay Calderon and Gal Yona introduced a knowledge profiling framework to examine factuality in Large Language Models (LLM), revealing that frontier LLM like Gemini3 and GPT-5 encode nearly all facts but struggle to recall many of them.
The framework uses WikiProfile, a benchmark of 2,150 Wikipedia-derived facts, to classify each fact into one of five knowledge profiles.
This approach provides a more informative diagnosis than question-level accuracy alone.
Why It Matters
⚡ Structural ImpactThe distinction between encoding and recall failures has significant implications for improving LLM reliability, as recall failures may be addressed through post-training and inference-time methods rather than solely relying on scaling model size or expanding data coverage. This understanding can lead to more targeted interventions to enhance factuality in LLM.
Multi-Vector Implications
- TECHNICALLLM may benefit from recall-enhancing techniques such as chain-of-thought prompting and thinking-optimized models
- MARKETImproved factuality in LLM can increase trust and adoption in applications requiring reliable information
- GOVERNANCEDevelopers should prioritize recall-focused evaluation metrics and methods to address factuality limitations in LLM
Strategic Outlook
🔭 12-18M HorizonOver the next 12-18 months, we can expect to see increased focus on developing and integrating recall-enhancing techniques into LLM architectures, potentially leading to significant improvements in factuality and reliability, with Google Research and other leaders in the field driving these advancements
Referenced Coverage & Sources
Read the full coverage below for original reporting, technical benchmarks, and complete primary source details.
How Amtrak Is Building the Data Backbone for Its Largest Transformation in Over 50 Years
"Every new trainset is a data-generating asset.
Rogue AI Agents Aren't Evil. They're Just Eager to Please
AI agents that break free and hack into other systems are only trying to make us happy.
Four of Five Enterprises That Secured AI Agent Identities Still Can't Contain One That Goes Rogue
Visa's president of technology, Rajat Taneja, walked the VB Transform 2026 audience through aiming Anthropic's Mythos at Visa's own payment network .
Twitch Streamers Can Now Opt Out From Training Amazon's AI
Twitch users can now opt out of allowing their content to be used to train Amazon's generative AI models.
Generative AI
Generative AI refers to algorithms and models designed to generate new, original content, including text, images, music, code, or video. Popular architectures like Transformers, GANs, and Diffusion models serve as the engines powering generative AI platforms.
Recall
Recall (Sensitivity or True Positive Rate) is a classification evaluation metric measuring the fraction of actual positive examples that the model correctly identified, calculated as true positives divided by all actual positives.
Explore technical glossaries, weekly market briefings, and editorial research articles related to this story:
Get top 5 high-signal AI news, venture funding rounds, and research papers auto-routed to dedicated channels every 3 hours.