Research Desk

How a High School Student's Algae Breakthrough Could Revolutionize Altitude Sensing

A 17-year-old high school student has successfully turned common algae into a biological altimeter that reached the stratosphere. Andrew's StratoSpore project combines spectral sensing with machine learning to measure altitude through algae fluorescence???a world first that could transform how we mo...

Read Full Article
Spoken Function Calling Rewrites Voice AI's Semantic Rules

Spoken Function Calling Rewrites Voice AI's Semantic Rules

A new arXiv preprint redefines how large audio language models should parse speech, shifting from rigid intent taxonomies to flexible function-calling contracts. The approach threatens established SLU vendors while handing open-weight LALM developers a practical path to open-domain voice automation.

WorldCup Arena: Live Tournament Exposes LLM Forecasting Lies

WorldCup Arena: Live Tournament Exposes LLM Forecasting Lies

The WorldCup Arena paper introduces a prospective benchmark design that eliminates memorization and contamination by asking models to forecast live sports outcomes before they happen. The findings reveal stark performance differences between frontier models and raise serious questions about the validity of all retrospective evaluations.

TurnSight Kills Trajectory-Level RL for Tool Agents

TurnSight Kills Trajectory-Level RL for Tool Agents

TurnSight replaces coarse trajectory-level rewards with dense, turn-level supervision derived from privileged context, promising faster convergence and better final performance in tool-use reasoning. This article breaks down what changed, who benefits, and what engineering tradeoffs matter.

ParVL Breaks MLLM Scaling: Fixed Compute Splits Are Dead

ParVL Breaks MLLM Scaling: Fixed Compute Splits Are Dead

ParVL introduces expandable compute allocation for multimodal LLMs, breaking the rigid vision-language compute split that limits task-specific optimization. The framework promises to reduce both memory and latency overhead compared to parameter or sequential scaling approaches.

GDPevo: The Benchmark That Exposes Fake Agent Evolution

GDPevo: The Benchmark That Exposes Fake Agent Evolution

GDPevo is a new evolution-native benchmark for AI agents grounded in GDP-related enterprise tasks. It targets the three core flaws of existing benchmarks: limited economic coverage, unisolated training effects, and data contamination, forcing a redefinition of what 'agent improvement' actually means.

HP Causes Expose Feature Attribution Flaws in Structured Inputs

HP Causes Expose Feature Attribution Flaws in Structured Inputs

The paper from arXiv (2608.03772v1) demonstrates that current explanation methods break under structured inputs. HP actual causes over SCMs fix this, but the approach demands a fundamental re-architecture of how AI observability platforms like Fiddler and Arize compute and present explanations.

Append the next batch without leaving this page.