On the estimation and validity of AI time horizons—a statistical look at the METR plot
BrickBench: Evaluating Agentic Brick Design
Ecology of AI Agents: Collaboration Creates a Population Threshold for Takeoff
Searching for “Harmful Refusal”: A Psychometric Audit of an AI Safety Benchmark
HRIL: Learning Multimodal Synergy via Higher-Order Tensor Modeling
GeoReform: Reflective Formalization Evolution for Multimodal Geometry Problem Solving
OnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal Transport
HANS: A Handwritten Answer Sheet Dataset for Noisy Hybrid Document Parsing
Cited but Not Consulted: A Counterfactual Audit of Legal Chain-of-Thought Faithfulness
Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict
Overcoming Prior Barriers: Supervised Fine-Tuning under Long-Tail Distribution
Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition
Prior or Feedback? What an LLM Uses When Adapting Neural Operators
Verdict Without the Rule: Diagnosing and Auditing Regulatory Rule Sensitivity in LLM Compliance Systems
A Structural Theory of Cognitive Representation and Problem Solving,Contexts, Invariance, and the Knowledge Space
Looking Inside LLMs: Small-World Connectivity as a Signature of Reasoning Performance
Learning Probabilistic Logic Programs with Functional Gradient Guided Language Models
One Word Opens the Gate: The Option-Channel Attack on Typed Decision Models as Agent Guardrails
When Has a Bayesian Neural Network Sampled Enough? Adaptive Inference Time with Statistical Guarantees
Recursive Self-Improvement through Multi-Agent Self-Supervision