A 2026 preprint reports faster AI-agent training updates, lower memory use and stronger scaling for psRL across four benchmark workloads.