AI Failure Study Reports Higher Repair Rates With Targeted Replay
A preprint reports higher failure reproduction and repair rates with selective replay than full reruns, while noting that benchmark results may not generalize.
A preprint reports higher failure reproduction and repair rates with selective replay than full reruns, while noting that benchmark results may not generalize.
A preprint evaluates TAU-Agent on traffic-video benchmarks, with second place on Track 3, fifth on PSI-VQA and 12th on FETV across leaderboard evaluations.
A preprint finds benchmark overlap can shift physical-AI model rankings, while four tests retain 78.5% of the full suiteās measured utility.
A preprint models neutron-star vortices and magnetic fluxtubes, reporting finite bouquets and type-1.5-like clustering without phase-gradient entrainment.
An arXiv preprint models charge-density-wave lock-in and finds that softer lattices absorb more mismatch, a prediction awaiting experimental tests.
A preprint compares three pruning methods and finds SparseGPT and WANDA keep fixed sparse-autoencoder metrics closer to dense-model results.
A 2026 arXiv preprint identifies fixed-evaluation quasimap stacks for Nakajima varieties with an explicit critical-locus construction.
A H.E.S.S. analysis finds a curved Galactic Center gamma-ray spectrum and places the preferred cosmic-ray injection site near Sgr A*.
A mathematics preprint links cyclic covers with mixed Hodge modules, singularity criteria, log canonical thresholds, and vanishing theorems.
A preprint reports faster CTPWA particle-physics fits on synthetic data, with the largest benchmark avoiding the conventional pipeline's memory failure.
A preprint finds opposing position and momentum trends, near-constant total entropy and more Gaussian-like ground-state behavior in a double well.
A preprint tests an object-centric model on KITTI-MOT, reporting top reconstruction scores but mixed results when the future camera is predicted.