alphaXiv connects papers, researchers, and organizations, grounding the answer in the underlying work.
Terence TaoTerence Tao's "Mathematics in the Age of AI" explores how the mathematical community should respond to advanced AI capabilities, arguing this era presents a "crisis of values and practices" analogous to the early 20th-century foundational crisis. The paper examines implicit mathematical goals such as verification, exposition, community acceptance, and canonicalization, positing they risk diverging from mere solution generation due to AI-driven optimization, and asserts that human understanding remains paramount.
An AI agent, Claude, autonomously conducts *de novo* protein binder design campaigns by orchestrating open-source tools and managing all design decisions from target research to final selection. The system achieved a 26.8% overall binder hit rate across 15 targets, with 90 designs exhibiting sub-10 nM affinity, and its performance matched or exceeded human expert teams in open competitions.
Recirculation introduces a training-free, inference-time architectural modification to large language models, enabling deep-layer contextualized representations to enrich shallower layers for improved state tracking. This method consistently reduced perplexity across various models and datasets, and enhanced performance on instruction following, contextualization, and reasoning tasks.
Managing Partner @ AI Aspire, Managing General Partner @ AI Fund, Executive Chairman @ LandingAI, Founder @ DeepLearning.AI, Adjunct Professor, CS @ Stanford University, Chairman and Co-Founder @ Coursera
Co-Founder @ Physical Intelligence, Assistant Professor, CS and EE @ Stanford University
Agentic ESOpt introduces an Evolution Strategies-based framework for full-parameter fine-tuning of long-horizon LLM agents, requiring only inference-level GPU memory. The method demonstrates superior performance over gradient-based reinforcement learning on complex tasks and facilitates optimization of larger language models and co-evolution with prompt-space techniques.
HarnessEval-W introduces an agentic evaluation framework for interactive visual world models, generating transparent "evidence trees" to explain model performance beyond scalar scores. The system aligns strongly with human preferences and provides fine-grained diagnoses for eight distinct evaluation axes, outperforming existing benchmarks in discriminative power.
Hydra-0 introduces "action flow," a kinematically grounded image-plane motion representation, to create a generalist world model capable of learning from and controlling diverse robot embodiments. The approach improves prediction fidelity across various datasets and enables both open-loop policy evaluation with high correlation to success rates and inverse control from desired object motion without task-specific robot demonstrations.
Microsoft researchers introduce Agent Lightning v1.0, a lightweight framework addressing the fundamental challenges of "harnessed agentic RL" where deploy-time agent harnesses are integrated into RL training. The framework systematically characterizes and provides solutions for issues like retokenization, advantage calculation, and loss normalization, demonstrating significant performance improvements on search, instruction-following, and coding agents, including a 14.6 percentage point gain on SWE-bench Verified.
A novel benchmark, ASI-Bench, evaluates AI systems' capacity for autonomous scientific research across 11 domains by progressively reducing human methodological guidance. Findings indicate current AI systems achieve an average score of 26.62 when independently determining research methods, compared to 50.91 with full guidance, suggesting that operationalizing scientific methods presents a significant challenge.
Luma AI researchers conducted a systematic scaling law study for text-to-image diffusion models, establishing that compute optimality occurs at approximately 200 image tokens per parameter, a 10x increase compared to large language models. The study also demonstrated that diffusion models exhibit scaling collapse and are robust to overtraining, providing empirically derived guidelines for efficient training.
Researchers from Google DeepMind, Carnegie Mellon, MIT, and Columbia University reduced the upper bound on the matrix multiplication exponent ω from 2.371339 to 2.371177 by employing modern gradient-based optimization and AI-driven algorithmic refinement with AlphaEvolve, enabling computation at an unprecedented scale.
Ali BehrouzProteus introduces incremental memory activation for long-context sequence modeling, which adaptively expands effective memory capacity during processing. This mechanism consistently improved average downstream accuracy and lowered perplexity across four memory-based model architectures, exhibiting enhanced robustness and length extrapolation, particularly at extended context lengths up to 16K.
Distinguished Scientist, Director of Efficient AI Research @ NVIDIA, Associate Professor, EECS @ Massachusetts Institute of Technology
The Dieter Schwarz Foundation Professor, CS & Senior Fellow, HAI @ Stanford University, Distinguished Scientist, Language and Cognition Research @ NVIDIA

Associate Professor (Provost’s Chair in AI) @ Nanyang Technological University, Previously Research Fellow @ The Chinese University of Hong Kong
GenRec, a multi-view flow matching model from ETH Zürich and Google, introduces an explicit architectural and supervisory split for reconstruction and generation in novel view synthesis. This approach achieves superior fidelity in observed regions and better perceptual quality in unobserved regions, outperforming existing baselines on various datasets with significantly faster inference times.
A "latent-to-pixel" adaptation strategy from Alibaba Token Hub enables the training of pixel-space text-to-image diffusion models that achieve quality comparable to or surpassing latent-space models. This approach demonstrated 3.18x to 4.75x end-to-end inference speedups over latent-space counterparts and a 100.6x speedup over the original 100 NFE latent-space pipeline.
τ0-VLA is a hierarchical robot foundation model that employs world-model-guided test-time computation for high-level subtask generation. This approach dynamically allocates computational resources to evaluate hypothetical subtask outcomes, leading to improved task success rates on long-horizon manipulation tasks and enhanced robustness in out-of-domain scenarios.
Peking University researchers introduced a framework for agentic 3D content creation through the joint design of an agent-centric Domain-Specific Language (aDSL) and a role-specialized multi-agent system. This approach consistently generated 3D models with improved robustness, control, and adherence to user intent across text-to-shape and image-to-shape tasks.
Meta researchers introduce MoE-ViE, a Mixture-of-Experts Vision Encoder that integrates a fine-grained MoE architecture, a magnitude-aware load balancing strategy, and a specialized Triton kernel to achieve state-of-the-art zero-shot performance on image and video benchmarks with enhanced efficiency. The largest model, MoE-ViE-H, matched or exceeded the performance of a 1.7 times larger dense encoder while demonstrating over 2.5 times faster inference compared to vanilla MoE implementations.
TAMP-Nav, a framework from the OmniAI Group of ZJU ACES Lab, introduces a "visual pointer" action formulation and an adaptive reasoning-on-demand mechanism with a lightweight memory system to enable efficient VLM-based embodied navigation. The framework attained state-of-the-art success rates on R2R-CE (66.2%) and RxR-CE (65.7%) while reducing inference latency to 16.58 seconds per task and demonstrating zero-shot real-world transferability (60% SR).
We study subcategories of the module category of a finite-dimensional algebra that are closed under quotients or submodules. We propose the quotient--submodule equidistribution conjecture: over a representation-finite algebra, the number of quotient-closed subcategories of size i is equal to that of submodule-closed subcategories of size i for every i, where size is the number of indecomposable modules in the subcategory. We prove the following cases of the conjecture: (1) the five smallest and five largest values of i, hence all algebras with at most nine indecomposable modules; (2) Nakayama algebras; (3) algebras with radical square zero; and (4) representation-directed algebras. In the last case, we classify quotient-closed subcategories by a lower Bruhat interval in a Coxeter group constructed from the Auslander--Reiten quiver, extending the classification of Oppermann--Reiten--Thomas for Dynkin quivers. We also prove that quotient closure defines a finitary convex geometry and classify the functorially finite quotient-closed subcategories by Gen-minimal modules.
Applying statistical mechanics principles to multi-agent large language model systems, research from Stanford University and UC Santa Barbara reveals that collective AI agent behaviors adhere to predictable dynamical laws. An Ising-like model, incorporating various interaction couplings, accurately forecasts individual opinion transitions (75 86% balanced accuracy) and reproduces observed group dynamics, elucidating the mechanisms behind conviction buildup, consensus, and truth-seeking or bias amplification.
GigaBrain-0.7 introduces a three-system architecture for embodied foundation models, coordinating understanding, prediction, and action with extensive heterogeneous data pretraining. It achieves improved generalization across diverse robot embodiments and tasks, demonstrating enhanced zero-shot capabilities, language-conditioned instruction following, and task success rates.