Papers
Deep, long-form breakdowns of recent research in AI, machine learning, security & programming languages.
A Zeroth-Order Paradigm for LLM Preference Alignment
The Likelihood Displacement Problem in Preference Alignment Training large language models to follow human preferences has become a core bottleneck in shipping reliable AI systems. Reinforcement learning from human feedback (RLHF)...
by Peter Chen, Xi Chen, Wotao Yin, Tianyi Lin
Read the deep dive
AgentLSD: Evaluating AI Security Agents Under Adversarial Task Contamination
by Matteo Golinelli, Idilio Drago, Matteo Boffa, Francesco Bergadano, Bruno Crispo
Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation
by Guanhua Ji, Tianyu Li, Dayoon Suh, Yuqian Zhang, Boyan Zhang, Nadia Figueroa
Exponential Hardness of Off-Policy Evaluation under History-Dependent Logging
by Pranaya Jajoo
Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments
by João Meneses dos Santos, Arlindo L. Oliveira
Flag Game: A Toy Model for Mechanistic Swarm Interpretability
by Elizabeth Pavlova, Hidenori Tanaka
Analog Pin Directionality as an Exfiltration Attack Surface in Mixed-Signal ICs
by Ramana Ranganatham, Chirag Adiga, Michael Zuzak, Tejasvi Das
rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference
by Kaijun Zhou, Zhiyang Li, Le Chen, Jinyu Gu
Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations
by Leon Bergen, Usha Bhalla, Andrew Lee, Barak Widawsky, Linas Nasvytis, Connor Watts, Siddharth Boppana, Sidharth Baskaran, Dron Hazra, Michael Byun, Atticus Geiger, Owen Lewis, Matthew Kowal, Vasudev Shyam, Thomas Fel, Thomas McGrath, Ekdeep Singh Lubana, Jack Merullo
Characterizing Network Centralization and Observability in the Remote MCP Ecosystem
by Muhammad Abdullah Sohail
Evidence-Grounded Agentic Formulation Development in an Autonomous Laboratory
by Michael M. Craig, Riley J. Hickman, Yingshan Ma, Rémi Piché-Taillefer, Christine Allen, Pauric Bannigan