TopheadlinesTopheadlines
    What's Hot

    Be careful who’s selling you that high-end AMD Radeon GPU, especially on Amazon

    December 2, 2023

    These are the most popular Science News stories of 2023

    January 2, 2024

    iOS 18.3.2 Update Coming Soon for iPhones

    March 10, 2025
    Facebook Twitter Instagram
    Trending
    • Simple blood markers that could detect ALS years before diagnosis
    • WNBA’s furious trans debate deepens as Seattle star takes bitter swipe at Sophie Cunningham… and protestors gather to support Indiana Fever icon
    • Cookies sold at major grocery store recalled over fears of life-threatening reaction
    • UNLV basketball program ‘not for sale’: Runnin’ Rebels coach Josh Pastner walks back wild claim
    • When Suzie hit her 40s, she put her thinning hair, irritability, wired brain and palpitations down to the perimenopause. Then she started to mysteriously lose weight…
    • Oil tumbles below $90 after pause in fighting between the US and Iran but experts still fear rates hike
    • Tour de France 2026: Pogacar makes history, cycling steps up anti-doping, and popularity grows
    • Trendy meat diet sparks explosion of crippling medieval disease in young people
    Facebook Twitter Instagram
    TopheadlinesTopheadlines
    • Latest News

      Inside Epstein’s Zorro Ranch ‘where paedophile looked to carry out human experiments and create super-race breeding facility’ and ‘buried girls who were strangled during sex’

      February 24, 2026

      Cheating Utah ‘grief author’ heard sobbing down phone to 911 operator after allegedly murdering husband with poisoned Moscow mule

      February 23, 2026

      Victim reveals how she was left with flesh-eating disease after GP did not see her face-to-face then fled to India

      February 23, 2026

      ‘Andrew, the prince of darkness’: Global media mark ‘the end of privilege’ for Mountbatten-Windsor and gloat that the royal is ‘at rock bottom’

      February 22, 2026

      Melania Trump stuns in silver pants as she arrives at controversial Governor’s Dinner with husband Donald as dozens threaten boycott after president’s turbulent week

      February 22, 2026
    • Politics

      Law enforcement says eight killed by avalanche in California mountains | Weather News

      February 18, 2026

      Bangladesh PM-to-be and lawmakers sworn into parliament | Sheikh Hasina

      February 17, 2026

      Hillary Clinton gets in testy exchange with European leader over Trump

      February 16, 2026

      At least 11 Palestinians killed in Israeli attacks across Gaza | Gaza News

      February 15, 2026

      DOJ sends letter to Congress with list of people named in Epstein files, including Trump: Report

      February 15, 2026
    • Tech

      Apple’s AI Wearables Expected to Lean Heavily on Visual Intelligence

      February 23, 2026

      What to Know About At-Home STI Tests: Pros, Cons, and Recommendations (2026)

      February 23, 2026

      10 Clever Home Appliance Innovations You’ll See in 2026

      February 22, 2026

      Runlayer is now offering secure OpenClaw agentic capabilities for large enterprises

      February 22, 2026

      Tesla's "cheaper" Cybertruck arrives at $59,990, still far from the $40K promise

      February 21, 2026
    • Business

      Oil tumbles below $90 after pause in fighting between the US and Iran but experts still fear rates hike

      July 27, 2026

      Chancellor urged to rule out pensions raid to stop repeat of damaging speculation under Reeves and back ‘people who do the right thing’

      July 24, 2026

      Does the risk of ruinous care bills deter YOU from spending or gifting money to beat inheritance tax?

      July 23, 2026

      This is the one step Andy Burnham could take now to make us all richer and happier – and it wouldn’t cost him a penny: RACHEL RICKARD STRAUS

      July 21, 2026

      Why will it take half a year to register our property purchase on the Land Registry and should we be concerned?

      July 17, 2026
    • Sports

      WNBA’s furious trans debate deepens as Seattle star takes bitter swipe at Sophie Cunningham… and protestors gather to support Indiana Fever icon

      July 29, 2026

      UNLV basketball program ‘not for sale’: Runnin’ Rebels coach Josh Pastner walks back wild claim

      July 28, 2026

      Tour de France 2026: Pogacar makes history, cycling steps up anti-doping, and popularity grows

      July 27, 2026

      The McCanns cheer on Madeleine’s brother Sean as he swims in Commonwealth Games nearly two decades after sister disappeared

      July 25, 2026

      Rockies vs. Brewers MLB picks: Keep riding same parlay in lone matinee Friday

      July 24, 2026
    • Health

      Simple blood markers that could detect ALS years before diagnosis

      July 29, 2026

      Cookies sold at major grocery store recalled over fears of life-threatening reaction

      July 28, 2026

      When Suzie hit her 40s, she put her thinning hair, irritability, wired brain and palpitations down to the perimenopause. Then she started to mysteriously lose weight…

      July 27, 2026

      Trendy meat diet sparks explosion of crippling medieval disease in young people

      July 26, 2026

      Scientists find potential cause of embarrassing condition that causes excessive sweating in 15 million people… paving way for new treatments

      July 25, 2026
    • Science

      Ever wonder where our math symbols came from? Here are their stories

      February 19, 2026

      Some dog breeds carry a higher risk of breathing problems

      February 19, 2026

      New study links early smartphone ownership to health risks

      February 18, 2026

      The Story of Stories traces the arc of storytelling across human history

      February 17, 2026

      Listen to the crackle of ‘mini-lightning’ on Mars

      February 16, 2026
    • Entertainment

      Paul McCartney and Wings Exhibit Set at Rock & Roll Hall of Fame

      February 18, 2026

      Warner Bros Latest Hollywood Studio To Warn Seedance Over AI Infringement

      February 18, 2026

      Cardi B on Stefon Diggs Relationship Status

      February 17, 2026

      Jelly Roll Will Receive Country Radio’s Humanitarian Award

      February 17, 2026

      X Down For Thousands In U.S. And UK

      February 16, 2026
    TopheadlinesTopheadlines
    Home»Tech»AI agent benchmarks are misleading, study warns
    Tech

    AI agent benchmarks are misleading, study warns

    July 7, 20247 Mins Read
    Facebook Twitter Pinterest LinkedIn WhatsApp Reddit Tumblr Email
    AI agent benchmarks are misleading, study warns
    Share
    Facebook Twitter LinkedIn Pinterest Email

    We want to hear from you! Take our quick AI survey and share your insights on the current state of AI, how you’re implementing it, and what you expect to see in the future. Learn More


    AI agents are becoming a promising new research direction with potential applications in the real world. These agents use foundation models such as large language models (LLMs) and vision language models (VLMs) to take natural language instructions and pursue complex goals autonomously or semi-autonomously. AI agents can use various tools such as browsers, search engines and code compilers to verify their actions and reason about their goals. 

    However, a recent analysis by researchers at Princeton University has revealed several shortcomings in current agent benchmarks and evaluation practices that hinder their usefulness in real-world applications.

    Their findings highlight that agent benchmarking comes with distinct challenges, and we can’t evaluate agents in the same way that we benchmark foundation models.

    Cost vs accuracy trade-off

    One major issue the researchers highlight in their study is the lack of cost control in agent evaluations. AI agents can be much more expensive to run than a single model call, as they often rely on stochastic language models that can produce different results when given the same query multiple times. 


    Countdown to VB Transform 2024

    Join enterprise leaders in San Francisco from July 9 to 11 for our flagship AI event. Connect with peers, explore the opportunities and challenges of Generative AI, and learn how to integrate AI applications into your industry. Register Now


    To increase accuracy, some agentic systems generate several responses and use mechanisms like voting or external verification tools to choose the best answer. Sometimes sampling hundreds or thousands of responses can increase the agent’s accuracy. While this approach can improve performance, it comes at a significant computational cost. Inference costs are not always a problem in research settings, where the goal is to maximize accuracy.

    However, in practical applications, there is a limit to the budget available for each query, making it crucial for agent evaluations to be cost-controlled. Failing to do so may encourage researchers to develop extremely costly agents simply to top the leaderboard. The Princeton researchers propose visualizing evaluation results as a Pareto curve of accuracy and inference cost and using techniques that jointly optimize the agent for these two metrics.

    The researchers evaluated accuracy-cost tradeoffs of different prompting techniques and agentic patterns introduced in different papers.

    “For substantially similar accuracy, the cost can differ by almost two orders of magnitude,” the researchers write. “Yet, the cost of running these agents isn’t a top-line metric reported in any of these papers.”

    The researchers argue that optimizing for both metrics can lead to “agents that cost less while maintaining accuracy.” Joint optimization can also enable researchers and developers to trade off the fixed and variable costs of running an agent. For example, they can spend more on optimizing the agent’s design but reduce the variable cost by using fewer in-context learning examples in the agent’s prompt.

    The researchers tested joint optimization on HotpotQA, a popular question-answering benchmark. Their results show that joint optimization formulation provides a way to strike an optimal balance between accuracy and inference costs.

    “Useful agent evaluations must control for cost—even if we ultimately don’t care about cost and only about identifying innovative agent designs,” the researchers write. “Accuracy alone cannot identify progress because it can be improved by scientifically meaningless methods such as retrying.”

    Model development vs downstream applications

    Another issue the researchers highlight is the difference between evaluating models for research purposes and developing downstream applications. In research, accuracy is often the primary focus, with inference costs being largely ignored. However, when developing real-world applications on AI agents, inference costs play a crucial role in deciding which model and technique to use.

    Evaluating inference costs for AI agents is challenging. For example, different model providers can charge different amounts for the same model. Meanwhile, the costs of API calls are regularly changing and might vary based on developers’ decisions. For example, on some platforms, bulk API calls are charged differently. 

    The researchers created a website that adjusts model comparisons based on token pricing to address this issue. 

    They also conducted a case study on NovelQA, a benchmark for question-answering tasks on very long texts. They found that benchmarks meant for model evaluation can be misleading when used for downstream evaluation. For example, the original NovelQA study makes retrieval-augmented generation (RAG) look much worse than long-context models than it is in a real-world scenario. Their findings show that RAG and long-context models were roughly equally accurate, while long-context models are 20 times more expensive.

    Overfitting is a problem

    In learning new tasks, machine learning (ML) models often find shortcuts that allow them to score well on benchmarks. One prominent type of shortcut is “overfitting,” where the model finds ways to cheat on the benchmark tests and provides results that do not translate to the real world. The researchers found that overfitting is a serious problem for agent benchmarks, as they tend to be small, typically consisting of only a few hundred samples. This issue is more severe than data contamination in training foundation models, as knowledge of test samples can be directly programmed into the agent.

    To address this problem, the researchers suggest that benchmark developers should create and keep holdout test sets that are composed of examples that can’t be memorized during training and can only be solved through a proper understanding of the target task. In their analysis of 17 benchmarks, the researchers found that many lacked proper holdout datasets, allowing agents to take shortcuts, even unintentionally. 

    “Surprisingly, we find that many agent benchmarks do not include held-out test sets,” the researchers write. “In addition to creating a test set, benchmark developers should consider keeping it secret to prevent LLM contamination or agent overfitting.”

    They also that different types of holdout samples are needed based on the desired level of generality of the task that the agent accomplishes.

    “Benchmark developers must do their best to ensure that shortcuts are impossible,” the researchers write. “We view this as the responsibility of benchmark developers rather than agent developers, because designing benchmarks that don’t allow shortcuts is much easier than checking every single agent to see if it takes shortcuts.”

    The researchers tested WebArena, a benchmark that evaluates the performance of AI agents in solving problems with different websites. They found several shortcuts in the training datasets that allowed the agents to overfit to tasks in ways that would easily break with minor changes in the real world. For example, the agent could make assumptions about the structure of web addresses without considering that it might change in the future or that it would not work on different websites.

    These errors inflate accuracy estimates and lead to over-optimism about agent capabilities, the researchers warn.

    With AI agents being a new field, the research and developer communities have yet much to learn about how to test the limits of these new systems that might soon become an important part of everyday applications.

    “AI agent benchmarking is new and best practices haven’t yet been established, making it hard to distinguish genuine advances from hype,” the researchers write. “Our thesis is that agents are sufficiently different from models that benchmarking practices need to be rethought.”

    VB Daily

    Stay in the know! Get the latest news in your inbox daily

    By subscribing, you agree to VentureBeat’s Terms of Service.

    Thanks for subscribing. Check out more VB newsletters here.

    An error occured.



    Source link
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Telegram Email

    Related Posts

    Apple’s AI Wearables Expected to Lean Heavily on Visual Intelligence

    February 23, 2026

    What to Know About At-Home STI Tests: Pros, Cons, and Recommendations (2026)

    February 23, 2026

    10 Clever Home Appliance Innovations You’ll See in 2026

    February 22, 2026
    Top Posts

    Simple blood markers that could detect ALS years before diagnosis

    July 29, 2026

    Francis Ngannou claims he is Man United’s ‘favourite heavyweight boxer’ as he teases Tyson Fury

    July 31, 2023

    Turn Your Favorite Pet Photos Into a Pawfect Portrait for Just $20

    July 31, 2023

    Mauricio Diazgranados Is a Botanist in a Hurry

    July 31, 2023
    Don't Miss
    Health

    Covid alert! New ‘super-contagious Frankenstein’ variant has rocketed four-fold in just a month…experts warn it could be most infectious yet

    Health 4 Mins Read

    By JOHN ELY DEPUTY HEALTH EDITOR FOR MAILONLINE Published: 16:51, 3 July 2025 | Updated:…

    Emotional King Charles leads impeccable Remembrance Sunday two minutes silence with Prince William, Kate, Queen Camilla, Edward and teary-eyed Sophie

    November 9, 2025

    I’ve Carried at Least 4 Wallets a Day for a Year to Test the Best Minimalist Wallets. These Are My Favorites

    January 13, 2026

    Agencies roll out initiative making it easier to click unsubscribe button

    August 12, 2024
    Stay In Touch
    • Facebook
    • Twitter
    • Instagram
    About Us
    About Us

    Delivering timely and accurate news updates, our website keeps you informed on the latest events, politics, entertainment, and more. Dive into captivating articles crafted by our team of expert writers, providing a comprehensive view of the world. Stay ahead with our trusted news source.

    Facebook Twitter Instagram
    Our Picks

    LLM Siri: The Complete Guide to Apple’s AI Assistant Overhaul Coming in 2026

    August 22, 2025

    A sixth mass extinction? Not so fast, some scientists say

    September 5, 2025

    US Republicans back Trump on Venezuela amid faint MAGA dissent | US-Venezuela Tensions News

    January 4, 2026
    © 2026 Designed by TopHeadlineSpot.
    • Business
    • Latest News
    • Politics
    • Health
    • Entertainment
    • Sports
    • Science
    • Tech

    Type above and press Enter to search. Press Esc to cancel.