Tarski Attack Destroys LLM Truth Probes
Abel Jansma's Tarski attack reveals that LLM truth probes are not measuring truth but a geometric artifact. This forces a reassessment of interpretability methods and their use in AI safety.
A 17-year-old high school student has successfully turned common algae into a biological altimeter that reached the stratosphere. Andrew's StratoSpore project combines spectral sensing with machine learning to measure altitude through algae fluorescence???a world first that could transform how we mo...
Read Full Article →
Abel Jansma's Tarski attack reveals that LLM truth probes are not measuring truth but a geometric artifact. This forces a reassessment of interpretability methods and their use in AI safety.

Anthropic’s Claude Mythos Preview has found new attacks against reduced-round AES variants, showing that AI can now perform cryptanalysis traditionally reserved for human experts. The finding challenges long-held assumptions about encryption safety margins and signals a new era for security auditing.

A new analysis shows that the standard way of handling classifier-free guidance in on-policy diffusion distillation is mathematically under-identified, leading to training instability. The authors propose a branch-level distillation objective that fixes the flaw without added cost.

TiRex-2 introduces a recurrent architecture that matches Transformer performance on multivariate forecasting while being computationally cheaper and natively supporting streaming data. This could reshape the foundation model landscape for time series.

ArXiv is overhauling its governance and funding model to handle explosive AI-driven growth. The changes could secure its future or dilute its founding mission, with major implications for the entire open-access research ecosystem.

A new research paper introduces a large-scale dataset of human similarity judgments conditioned on free-form text aspects, exposing that current AI vision metrics fail to capture context-dependent similarity. This work will likely become a critical benchmark, forcing a re-evaluation of how perceptual metrics are designed.

A new study shows that a simple domain generalization method can beat complex AI at detecting image tampering from VLMs like ChatGPT and Gemini. The finding suggests that current forensic AI may be overengineered, but also raises questions about robustness against future models.

Flödb.ai's two-bit bloom filter claims to double accuracy at no memory cost. We analyze the evidence, the limits, and who wins if this holds up.

A new study shows AI advice makes people 3x less accurate but 2x more confident, raising urgent questions about AI deployment in high-stakes decisions. The findings challenge the entire premise of AI as an augmentation tool.

MeanFlowNFT adapts the forward-process RL framework of DiffusionNFT to average-velocity generators, enabling fast few-step sampling with preference alignment. While technically sound, the incremental nature of the contribution raises questions about its practical impact versus existing methods.

Hindcast demonstrates that current backtesting methods for LLM forecasters are invalid due to systematic data leakage, overstating accuracy by 30% or more. This forces a reckoning for every lab and platform that uses these benchmarks to claim forecasting competence.

A new framework from arXiv reveals that the forensic capabilities of watermarks in generative text follow a strict ladder, each rung costing more in sample length. The findings challenge the adequacy of current detection-only systems.
Append the next batch without leaving this page.