Close Menu
  • Academy
  • Events
  • Identity
  • International
  • Inventions
  • Startups
    • Sustainability
  • Tech
  • Spanish
What's Hot

Malicious browser extensions will infect 722 users across Latin America since early 2025

Trump officials vow to lift school separation orders

Should the government ban AI-generated humans to stop the collapse of social trust?

Facebook X (Twitter) Instagram
  • Home
  • About Us
  • Advertise with Us
  • Contact Us
  • DMCA
  • Privacy Policy
  • Terms & Conditions
  • User-Submitted Posts
Facebook X (Twitter) Instagram
Fyself News
  • Academy
  • Events
  • Identity
  • International
  • Inventions
  • Startups
    • Sustainability
  • Tech
  • Spanish
Fyself News
Home » The new AI model meta benchmark is a bit misleading
Startups

The new AI model meta benchmark is a bit misleading

userBy userApril 6, 2025No Comments2 Mins Read
Share Facebook Twitter Pinterest Telegram LinkedIn Tumblr Email Copy Link
Follow Us
Google News Flipboard
Share
Facebook Twitter LinkedIn Pinterest Email Copy Link

One of the new flagship AI model meta released on Saturday, Maverick ranks second in the LM Arena. This is a test in which a human evaluator compares the output of the model and selects preferences. However, the version of Maverick that Meta deployed in LM Arena appears to be different from the version widely available to developers.

As some AI researchers pointed out in X, Meta said that Maverick of LM Arena has announced that it is an “experimental chat version.” Meanwhile, the chart on the official Llama website reveals that Meta’s LM Arena test was conducted using “Llama 4 Maverick optimized for conversation.”

As I wrote before, for a variety of reasons, LM arena was not the most reliable measure of AI models’ performance. However, AI companies generally do not customize or tweak their models, or at least allow them to do so, in order to score better at LM Arena.

The problem with adjusting the model to its benchmark, withholding it, then releasing a “vanilla” variant of the same model is that it becomes difficult for developers to accurately predict the performance of the model in a given context. That’s also misleading. Ideally, the benchmark is as badly insufficient as it is – providing a snapshot of the advantages and disadvantages of a single model across a variety of tasks.

In fact, X researchers have observed significant differences in the behavior of publicly available Mavericks compared to models hosted at LM Arena. The LM Arena version seems to use a lot of emojis and provide a very long answer.

OK llama4 is a lol with def cooked.

– Nathan Lambert (@natolambert) April 6, 2025

For some reason, the Arena Lama 4 model uses more emojis

together. ai, it seems better: pic.twitter.com/f74odx4ztt

– Tech Dev Notes (@techdevnotes) April 6, 2025

For comments, we contacted Chatbot Arena with Meta, the organization that maintains LM Arena.




Source link

Follow on Google News Follow on Flipboard
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email Copy Link
Previous ArticleYemen’s Hootis says the latest US attacks kill at least four at Sanaa | News
Next Article Pennylane doubles its valuation as Alphabet VC Fund acquires shares
user
  • Website

Related Posts

Lawyers could face “severe” penalties for quotes generated by fake AI, UK courts warn

June 7, 2025

Review Week: Why Humanity’s Cut Access to Windsurf

June 7, 2025

Will Musk vs. Trump affect Xai’s $5 billion debt transaction?

June 7, 2025
Add A Comment
Leave A Reply Cancel Reply

Latest Posts

Malicious browser extensions will infect 722 users across Latin America since early 2025

Trump officials vow to lift school separation orders

Should the government ban AI-generated humans to stop the collapse of social trust?

Lawyers could face “severe” penalties for quotes generated by fake AI, UK courts warn

Trending Posts

Sana Yousaf, who was the Pakistani Tiktok star shot by gunmen? |Crime News

June 4, 2025

Trump says it’s difficult to make a deal with China’s xi’ amid trade disputes | Donald Trump News

June 4, 2025

Iraq’s Jewish Community Saves Forgotten Shrine Religious News

June 4, 2025

Subscribe to News

Subscribe to our newsletter and never miss our latest news

Please enable JavaScript in your browser to complete this form.
Loading

Welcome to Fyself News, your go-to platform for the latest in tech, startups, inventions, sustainability, and fintech! We are a passionate team of enthusiasts committed to bringing you timely, insightful, and accurate information on the most pressing developments across these industries. Whether you’re an entrepreneur, investor, or just someone curious about the future of technology and innovation, Fyself News has something for you.

Should the government ban AI-generated humans to stop the collapse of social trust?

AB will be released at Binance -Tech Startups

Top 10 Startups and Tech Funding News for the Weekly Ends June 6, 2025

Order openai to keep all chatgpt logs including deleted temporary chats, API requests

Facebook X (Twitter) Instagram Pinterest YouTube
  • Home
  • About Us
  • Advertise with Us
  • Contact Us
  • DMCA
  • Privacy Policy
  • Terms & Conditions
  • User-Submitted Posts
© 2025 news.fyself. Designed by by fyself.

Type above and press Enter to search. Press Esc to cancel.