ai Sep 17, 2026

A Zeroth-Order Paradigm for LLM Preference Alignment

The Likelihood Displacement Problem in Preference Alignment Training large language models to follow human preferences has become a core bottleneck in shipping reliable AI systems. Reinforcement learning from human feedback (RLHF)...

by Peter Chen, Xi Chen, Wotao Yin, Tianyi Lin Read the deep dive
cybersecurity Sep 17 AgentLSD: Evaluating AI Security Agents Under Adversarial Task Contamination by Matteo Golinelli, Idilio Drago, Matteo Boffa, Francesco Bergadano, Bruno Crispo ai Sep 17 Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation by Guanhua Ji, Tianyu Li, Dayoon Suh, Yuqian Zhang, Boyan Zhang, Nadia Figueroa machine learning Sep 17 Exponential Hardness of Off-Policy Evaluation under History-Dependent Logging by Pranaya Jajoo ai Sep 17 Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments by João Meneses dos Santos, Arlindo L. Oliveira ai Sep 17 Flag Game: A Toy Model for Mechanistic Swarm Interpretability by Elizabeth Pavlova, Hidenori Tanaka cybersecurity Sep 17 Analog Pin Directionality as an Exfiltration Attack Surface in Mixed-Signal ICs by Ramana Ranganatham, Chirag Adiga, Michael Zuzak, Tejasvi Das machine learning Sep 17 rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference by Kaijun Zhou, Zhiyang Li, Le Chen, Jinyu Gu ai Sep 17 Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations by Leon Bergen, Usha Bhalla, Andrew Lee, Barak Widawsky, Linas Nasvytis, Connor Watts, Siddharth Boppana, Sidharth Baskaran, Dron Hazra, Michael Byun, Atticus Geiger, Owen Lewis, Matthew Kowal, Vasudev Shyam, Thomas Fel, Thomas McGrath, Ekdeep Singh Lubana, Jack Merullo cybersecurity Sep 17 Characterizing Network Centralization and Observability in the Remote MCP Ecosystem by Muhammad Abdullah Sohail ai Sep 17 Evidence-Grounded Agentic Formulation Development in an Autonomous Laboratory by Michael M. Craig, Riley J. Hickman, Yingshan Ma, Rémi Piché-Taillefer, Christine Allen, Pauric Bannigan