Coverage hub
Research
Papers, benchmarks, training techniques, and measurable progress toward AGI.
MIT Report Finds AI Eroding Academic Engagement and Faculty-Student Trust
The committee advises against AI detection software and recommends course-specific policies over institute-wide bans.
ResearchDeepmind Researchers Propose Artificial Symbiotic Intelligence as Alternative to Singularity
The authors argue that intelligence is a social phenomenon, requiring governance of complex human-machine networks rather than isolated superintelligences.
ResearchResearchers Argue Large Language Models Lack Genuine Reasoning Capabilities
Unlike AlphaGo, which combined intuition with explicit search, current LLMs rely solely on next-token prediction.
ResearchOpenAI Cuts Ties With 3 Safety Researchers, WSJ Reports
The departures follow an internal investigation into the mishandling of sensitive company information shared with a third-party group.
ResearchAI Researchers Warn Superintelligence Is ‘Exactly as Dangerous as It Sounds’
Nonprofit Palisade Research released a series of interviews with current and former staff from major AI labs discussing existential risks.
ResearchHugging Face Launches Open TTS Leaderboard to Evaluate Multilingual Models
The new leaderboard uses objective metrics like word error rate and speed to rank open-source models in hours rather than weeks.
ResearchStudy Finds AI Access Makes People Unwilling to Say 'I Don't Know'
Research suggests that reliance on AI tools reduces users' comfort with admitting uncertainty.
ResearchStudy Finds Top AI Experts Underestimated Field's Pace of Progress
Surveyed experts missed specific 2024-2025 model benchmarks, with actual releases outpacing forecasts by 12-18 months.
ResearchMultiverse Computing Frames LLM Block Removal as Ising Optimization Problem
The new approach maps the selection of transformer blocks to a physics-based optimization model to streamline model pruning.
ResearchAnthropic Partners With Accenture on Embedded Evaluation
The collaboration focuses on integrating evaluation frameworks directly into Accenture's client delivery models to ensure AI reliability.
ResearchOpenAI Model Repeatedly Inserts Prompt Injections Into Its Own Notes, Researchers Say
Researchers report the unexpected behavior remains unexplained, describing it as weird rather than a confirmed system flaw.
ResearchTwo-Year University Study Finds Banning AI From Classrooms Leaves Students Worse Off
Research indicates that prohibiting artificial intelligence tools resulted in lower academic performance compared to classes that allowed their use.
ResearchWhy AI Food Images Often Look Unusual, Explained
Researchers point to training data imbalances and visual ambiguity as key factors in AI food rendering.
ResearchHugging Face Blog Details Fine-Tuning a 350M Model for Structured Outputs using GRPO
The approach uses TRL's IfStruct framework and achieves improvements in 100 optimization steps.
ResearchResearchers Fear Safety Disaster Ahead of OpenAI's Astra Release
External experts say internal safety monitoring at the company may be insufficient for a system as capable as Astra.
ResearchAllenAI Researchers Examine What LLM Benchmarks Actually Measure
The team introduces BenchMIRT, a framework designed to assess whether current evaluation methods capture meaningful reasoning.
ResearchLAION Releases Open Video Dataset With 10 Million Hours of Footage for AI Research
The nonprofit organization that previously created popular image datasets has expanded into video, offering the collection at no cost.
ResearchAI Benchmarks Have a Trust Problem and Google Wants to Fix It
The company is proposing new evaluation standards amid concerns about benchmark reliability.
ResearchGoogle DeepMind AI Co-Scientist Can Now Plan Experiments, Operate Lab Equipment and Write Papers
The system represents an expansion of the AI assistant's capabilities beyond prior versions.
ResearchGoogle DeepMind Pilots Double-Blind Evaluations for AI Systems
New methodology aims to reduce bias in how frontier AI models are assessed.
ResearchPew Study Confirms Sharp Rise of AI-Written Text on the Web Since ChatGPT's Launch
Researchers find proportion of AI-detectable content online increased significantly after late 2022.
ResearchStanford Study Finds AI Is Hitting Entry-Level Jobs Hardest
Workers aged 22 to 25 in AI-exposed occupations are now 19% below peers in less exposed fields, up from 13% last year.
ResearchStudy Suggests AI Could Lead Scientists to Produce More Work at Lower Quality
A research paper argues automation tools may increase output volume while reducing depth of analysis.
ResearchMeasuring Benchmark Optimization in Speech Recognition
Hugging Face researchers examine how automatic speech recognition models may be tuned to perform well on specific benchmarks.
ResearchResearchers Say OpenAI Revoked Their Access to Limited Cyber Program
External researchers report losing access to a cybersecurity-focused program, raising questions about transparency.
ResearchTerence Tao Says AI Could Trigger Mathematics' Biggest Crisis Since Gödel
The Fields Medalist warns that AI-generated proofs may outpace human verification and erode peer review standards.
ResearchA Third of Web Pages Published Since ChatGPT's Launch Show Signs of AI Authorship, Study Finds
The study analyzed pages published after November 2022 when OpenAI launched its popular chatbot.
ResearchOptima Targets AI Benchmarking Limitation by Letting Users Test Models Against Their Own Data
The platform allows developers to evaluate models against their own datasets rather than relying on standardized tests.
ResearchGoogle Research's AMIE Medical AI System Demonstrates Real-Time Clinical Video Consultation Capabilities in First-of-Its-Kind Study
The research system, designed for medical intelligence, shows the ability to conduct live clinical video consultations according to Google.
ResearchResearchers Find Unexpected Content Including Recipe References in ChatGPT's Internal Reasoning Traces
The investigation uncovered marinade recipes alongside credential-like strings persisting within ChatGPT's chain-of-thought traces.