Papers
Deep, long-form breakdowns of recent research in AI, machine learning, security & programming languages.
49 results on this page · clear filters
Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models
by Wonje Jeung, Sangyeon Yoon, Hyesoo Hong, Yoonjun Cho, Dongjae Jeon, Bumjun Kim, Jean Oh, Youngjae Yu, Albert No
Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence
by Urja Pawar, Rajitha Ramanayake, Nabeel Kemal, Ashwin Kandath, Owen O'Neill, Guillaume Bourgeon, Houssem Chatbri
Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models
by Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Roland C. Aydin, Christian Feiler
Propagation Model for SSC attacks: Why SBOM (tools) don't tell the whole truth
by Ljubica Grgic, Lazar Maksimovic, Pavel Laskov
CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents
by Haoting Shi, Wenhao Wang, Weicheng Fang, Yaozhong Liang, Tian Jin, Pengxiang Zhao, Guangyi Liu, Siheng Chen, Yanfeng Wang
When LLM Decompilers Recompile More and Preserve Less
by Chang Liu, Edward Raff, Kristopher Micinski
Design Docs Are All You Need: An AI-native Machine-Learning Performance Tool
by Samuel Kushnir, Kimia Noorbakhsh, Kavya Sreedhar, Liqun Cheng, Ming Liu, Parthasarathy Ranganathan, Mohammad Alizadeh, Fred Kjolstad, Suvinay Subramanian
Distill Globally, Adapt Locally: Reasoning Distillation and Product-Type Test-Time Training for Scalable Trade-Up Recommendation
by Siliang Liu, Mohammad Ghasemi, Sapan Patel, Amin Banitalebi-Dehkordi
MEOX: Compact Multimodal Mixture-of-Experts for Earth Observation
by Mohanad Albughdadi