TopheadlinesTopheadlines
    What's Hot

    Biden gives the Presidential Medal of Freedom to Michelle Yeoh, Nancy Pelosi, John Kerry, Mike Bloomberg and Al Gore … and slides in a dig at Trump

    May 4, 2024

    First Look at Messages via Satellite in iOS 18

    June 13, 2024

    The battery boom is reviving US graphite mining after decades of dormancy

    December 28, 2025
    Facebook Twitter Instagram
    Trending
    • Simple blood markers that could detect ALS years before diagnosis
    • WNBA’s furious trans debate deepens as Seattle star takes bitter swipe at Sophie Cunningham… and protestors gather to support Indiana Fever icon
    • Cookies sold at major grocery store recalled over fears of life-threatening reaction
    • UNLV basketball program ‘not for sale’: Runnin’ Rebels coach Josh Pastner walks back wild claim
    • When Suzie hit her 40s, she put her thinning hair, irritability, wired brain and palpitations down to the perimenopause. Then she started to mysteriously lose weight…
    • Oil tumbles below $90 after pause in fighting between the US and Iran but experts still fear rates hike
    • Tour de France 2026: Pogacar makes history, cycling steps up anti-doping, and popularity grows
    • Trendy meat diet sparks explosion of crippling medieval disease in young people
    Facebook Twitter Instagram
    TopheadlinesTopheadlines
    • Latest News

      Inside Epstein’s Zorro Ranch ‘where paedophile looked to carry out human experiments and create super-race breeding facility’ and ‘buried girls who were strangled during sex’

      February 24, 2026

      Cheating Utah ‘grief author’ heard sobbing down phone to 911 operator after allegedly murdering husband with poisoned Moscow mule

      February 23, 2026

      Victim reveals how she was left with flesh-eating disease after GP did not see her face-to-face then fled to India

      February 23, 2026

      ‘Andrew, the prince of darkness’: Global media mark ‘the end of privilege’ for Mountbatten-Windsor and gloat that the royal is ‘at rock bottom’

      February 22, 2026

      Melania Trump stuns in silver pants as she arrives at controversial Governor’s Dinner with husband Donald as dozens threaten boycott after president’s turbulent week

      February 22, 2026
    • Politics

      Law enforcement says eight killed by avalanche in California mountains | Weather News

      February 18, 2026

      Bangladesh PM-to-be and lawmakers sworn into parliament | Sheikh Hasina

      February 17, 2026

      Hillary Clinton gets in testy exchange with European leader over Trump

      February 16, 2026

      At least 11 Palestinians killed in Israeli attacks across Gaza | Gaza News

      February 15, 2026

      DOJ sends letter to Congress with list of people named in Epstein files, including Trump: Report

      February 15, 2026
    • Tech

      Apple’s AI Wearables Expected to Lean Heavily on Visual Intelligence

      February 23, 2026

      What to Know About At-Home STI Tests: Pros, Cons, and Recommendations (2026)

      February 23, 2026

      10 Clever Home Appliance Innovations You’ll See in 2026

      February 22, 2026

      Runlayer is now offering secure OpenClaw agentic capabilities for large enterprises

      February 22, 2026

      Tesla's "cheaper" Cybertruck arrives at $59,990, still far from the $40K promise

      February 21, 2026
    • Business

      Oil tumbles below $90 after pause in fighting between the US and Iran but experts still fear rates hike

      July 27, 2026

      Chancellor urged to rule out pensions raid to stop repeat of damaging speculation under Reeves and back ‘people who do the right thing’

      July 24, 2026

      Does the risk of ruinous care bills deter YOU from spending or gifting money to beat inheritance tax?

      July 23, 2026

      This is the one step Andy Burnham could take now to make us all richer and happier – and it wouldn’t cost him a penny: RACHEL RICKARD STRAUS

      July 21, 2026

      Why will it take half a year to register our property purchase on the Land Registry and should we be concerned?

      July 17, 2026
    • Sports

      WNBA’s furious trans debate deepens as Seattle star takes bitter swipe at Sophie Cunningham… and protestors gather to support Indiana Fever icon

      July 29, 2026

      UNLV basketball program ‘not for sale’: Runnin’ Rebels coach Josh Pastner walks back wild claim

      July 28, 2026

      Tour de France 2026: Pogacar makes history, cycling steps up anti-doping, and popularity grows

      July 27, 2026

      The McCanns cheer on Madeleine’s brother Sean as he swims in Commonwealth Games nearly two decades after sister disappeared

      July 25, 2026

      Rockies vs. Brewers MLB picks: Keep riding same parlay in lone matinee Friday

      July 24, 2026
    • Health

      Simple blood markers that could detect ALS years before diagnosis

      July 29, 2026

      Cookies sold at major grocery store recalled over fears of life-threatening reaction

      July 28, 2026

      When Suzie hit her 40s, she put her thinning hair, irritability, wired brain and palpitations down to the perimenopause. Then she started to mysteriously lose weight…

      July 27, 2026

      Trendy meat diet sparks explosion of crippling medieval disease in young people

      July 26, 2026

      Scientists find potential cause of embarrassing condition that causes excessive sweating in 15 million people… paving way for new treatments

      July 25, 2026
    • Science

      Ever wonder where our math symbols came from? Here are their stories

      February 19, 2026

      Some dog breeds carry a higher risk of breathing problems

      February 19, 2026

      New study links early smartphone ownership to health risks

      February 18, 2026

      The Story of Stories traces the arc of storytelling across human history

      February 17, 2026

      Listen to the crackle of ‘mini-lightning’ on Mars

      February 16, 2026
    • Entertainment

      Paul McCartney and Wings Exhibit Set at Rock & Roll Hall of Fame

      February 18, 2026

      Warner Bros Latest Hollywood Studio To Warn Seedance Over AI Infringement

      February 18, 2026

      Cardi B on Stefon Diggs Relationship Status

      February 17, 2026

      Jelly Roll Will Receive Country Radio’s Humanitarian Award

      February 17, 2026

      X Down For Thousands In U.S. And UK

      February 16, 2026
    TopheadlinesTopheadlines
    Home»Tech»DeepMind’s Michelangelo benchmark reveals limitations of long-context LLMs
    Tech

    DeepMind’s Michelangelo benchmark reveals limitations of long-context LLMs

    October 11, 20246 Mins Read
    Facebook Twitter Pinterest LinkedIn WhatsApp Reddit Tumblr Email
    DeepMind’s Michelangelo benchmark reveals limitations of long-context LLMs
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More


    Large language models (LLMs) with very long context windows have been making headlines lately. The ability to cram hundreds of thousands or even millions of tokens into a single prompt unlocks many possibilities for developers. 

    But how well do these long-context LLMs really understand and utilize the vast amounts of information they receive?

    Researchers at Google DeepMind have introduced Michelangelo, a new benchmark designed to evaluate the long-context reasoning capabilities of LLMs. Their findings, published in a new research paper, show that while current frontier models have progressed in retrieving information from large in-context data, they still struggle with tasks that require reasoning over the data structure.

    The need for better long-context benchmarks

    The emergence of LLMs with extremely long context windows, ranging from 128,000 to over 1 million tokens, has prompted researchers to develop new benchmarks to evaluate their capabilities. However, most of the focus has been on retrieval tasks, such as the popular “needle-in-a-haystack” evaluation, where the model is tasked with finding a specific piece of information within a large context.

    “Over time, models have grown considerably more capable in long context performance,” Kiran Vodrahalli, research scientist at Google DeepMind, told VentureBeat. “For instance, the popular needle-in-a-haystack evaluation for retrieval has now been well saturated up to extremely long context lengths. Thus, it has become important to determine whether the harder tasks models are capable of solving in short context regimes are also solvable at long ranges.”

    Retrieval tasks don’t necessarily reflect a model’s capacity for reasoning over the entire context. A model might be able to find a specific fact without understanding the relationships between different parts of the text. Meanwhile, existing benchmarks that evaluate a model’s ability to reason over long contexts have limitations.

    “It is easy to develop long reasoning evaluations which are solvable with a combination of only using retrieval and information stored in model weights, thus ‘short-circuiting’ the test of the model’s ability to use the long-context,” Vodrahalli said.

    Michelangelo

    To address the limitations of current benchmarks, the researchers introduced Michelangelo, a “minimal, synthetic, and unleaked long-context reasoning evaluation for large language models.” 

    Michelangelo is based on the analogy of a sculptor chiseling away irrelevant pieces of marble to reveal the underlying structure. The benchmark focuses on evaluating the model’s ability to understand the relationships and structure of the information within its context window, rather than simply retrieving isolated facts.

    The benchmark consists of three core tasks:

    Latent list: The model must process a long sequence of operations performed on a Python list, filter out irrelevant or redundant statements, and determine the final state of the list. “Latent List measures the ability of a model to track a latent data structure’s properties over the course of a stream of code instructions,” the researchers write.

    Multi-round co-reference resolution (MRCR): The model must produce parts of a long conversation between a user and an LLM. This requires the model to understand the structure of the conversation and resolve references to previous turns, even when the conversation contains confusing or distracting elements. “MRCR measures the model’s ability to understanding ordering in natural text, to distinguish between similar drafts of writing, and to reproduce a specified piece of previous context subject to adversarially difficult queries,” the researchers write.

    “I don’t know” (IDK): The model is given a long story and asked to answer multiple-choice questions about it. For some questions, the context does not contain the answer, and the model must be able to recognize the limits of its knowledge and respond with “I don’t know.” “IDK measures the model’s ability to understand whether it knows what it doesn’t know based on the presented context,” the researchers write.

    Latent Structure Queries

    The tasks in Michelangelo are based on a novel framework called Latent Structure Queries (LSQ). LSQ provides a general approach for designing long-context reasoning evaluations that can be extended to arbitrary lengths. It can also test the model’s understanding of implicit information as opposed to retrieving simple facts. LSQ relies on synthesizing test data to avoid the pitfalls of test data leaking into the training corpus.

    “By requiring the model to extract information from structures rather than values from keys (sculptures from marble rather than needles from haystacks), we can more deeply test language model context understanding beyond retrieval,” the researchers write.

    LSQ has three key differences from other approaches to evaluating long-context LLMs. First, it has been explicitly designed to avoid short-circuiting flaws in evaluations that go beyond retrieval tasks. Second, it specifies a methodology for increasing task complexity and context length independently. And finally, it is general enough to capture a large range of reasoning tasks. The three tests used in Michelangelo cover code interpretation and reasoning over loosely written text.

    “The goal is that long-context beyond-reasoning evaluations implemented by following LSQ will lead to fewer scenarios where a proposed evaluation reduces to solving a retrieval task,” Vodrahalli said.

    Evaluating frontier models on Michelangelo

    The researchers evaluated ten frontier LLMs on Michelangelo, including different variants of Gemini, GPT-4 and 4o, and Claude. They tested the models on contexts up to 1 million tokens. Gemini models performed best on MRCR, GPT models excelled on Latent List, and Claude 3.5 Sonnet achieved the highest scores on IDK.

    However, all models exhibited a significant drop in performance as the complexity of the reasoning tasks increased, suggesting that even with very long context windows, current LLMs still have room to improve in their ability to reason over large amounts of information.

    Frontier LLMs struggle with reasoning on long-context windows (source: arxiv)

    “Frontier models have room to improve on all of the beyond-retrieval reasoning primitives (Latent List, MRCR, IDK) that we investigate in Michelangelo,” Vodrahalli said. “Different frontier models have different strengths and weaknesses – each class performs well on different context ranges and on different tasks. What does seem to be universal across models is the initial drop in performance on long reasoning tasks.”

    The Michelangelo evaluations capture basic primitives necessary for long-context reasoning and the findings can have important implications for enterprise applications. For example, in real-world applications where the model can’t rely on its pretraining knowledge and must perform multi-hop reasoning over many disparate locations in very long contexts, Vodrahalli expects performance to drop as the context length grows.

    “This is particularly true if the documents have a lot of information that is irrelevant to the task at hand, making it hard for a model to easily immediately distinguish which information is relevant or not,” Vodrahalli said. “It is also likely that models will continue to perform well on tasks where all of the relevant information to answer a question is located in one general spot in the document.”

    The researchers will continue to add more evaluations to Michelangelo and hope to make them directly available so that other researchers can test their models on them.

    VB Daily

    Stay in the know! Get the latest news in your inbox daily

    By subscribing, you agree to VentureBeat’s Terms of Service.

    Thanks for subscribing. Check out more VB newsletters here.

    An error occured.



    Source link
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Telegram Email

    Related Posts

    Apple’s AI Wearables Expected to Lean Heavily on Visual Intelligence

    February 23, 2026

    What to Know About At-Home STI Tests: Pros, Cons, and Recommendations (2026)

    February 23, 2026

    10 Clever Home Appliance Innovations You’ll See in 2026

    February 22, 2026
    Top Posts

    Simple blood markers that could detect ALS years before diagnosis

    July 29, 2026

    Francis Ngannou claims he is Man United’s ‘favourite heavyweight boxer’ as he teases Tyson Fury

    July 31, 2023

    Turn Your Favorite Pet Photos Into a Pawfect Portrait for Just $20

    July 31, 2023

    Mauricio Diazgranados Is a Botanist in a Hurry

    July 31, 2023
    Don't Miss
    Sports

    UCLA’s Lauren Betts has near-perfect performance, reaches 1,000 career points in top-10 win over Maryland

    Sports 2 Mins Read

    Getty Images UCLA star Lauren Betts registered a career-high 33 points in an 82-67 road…

    Dua Lipa turns heads in sparkling red gown as she joins Claudia Schiffer and Maura Higgins at star-studded premiere of Argylle in Leicester Square

    January 24, 2024

    New embedding model leaderboard shakeup: Google takes #1 while Alibaba’s open source alternative closes gap

    July 19, 2025

    Antidepressants: What to Know About Uses and Side Effects

    April 25, 2024
    Stay In Touch
    • Facebook
    • Twitter
    • Instagram
    About Us
    About Us

    Delivering timely and accurate news updates, our website keeps you informed on the latest events, politics, entertainment, and more. Dive into captivating articles crafted by our team of expert writers, providing a comprehensive view of the world. Stay ahead with our trusted news source.

    Facebook Twitter Instagram
    Our Picks

    Diddy Denied $50 Million Bond Proposal to Get Out of Jail After Arrest

    September 17, 2024

    Bellevue Hospital Bariatric Surgery Program Is Under NY State Scrutiny

    December 21, 2023

    Peter Seidler, Big-Spending San Diego Padres Owner, Dies at 63

    December 28, 2023
    © 2026 Designed by TopHeadlineSpot.
    • Business
    • Latest News
    • Politics
    • Health
    • Entertainment
    • Sports
    • Science
    • Tech

    Type above and press Enter to search. Press Esc to cancel.