Papers
Deep, long-form breakdowns of recent research in AI, machine learning, security & programming languages.
7 results on this page · clear filters
Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation
by Daniel P. Jeong, Charles Q. Li, Hossein Hosseiny, Nitya M. Bhalla, Fatma Uyar Morency, Pradeep Ravikumar, Zachary C. Lipton, Michael Oberst
Right Tool, Right Job: Native-Language Evaluation, Tokenizer Sensitivity, and Methodological Findings from a French-Only BabyLM
by Adam Zachary Wasserman, David Beauchemin
Verifiable by Construction: Claim-Level Evaluation of Verbatim Citation in Clinical Question Answering
by Jiashuo Zhang, Yuling Chen, Yvonne Commodore-Mensah, Michael Oberst
SNAP3D: Physically Grounded 3D Parts for Assembly from a Single Image
by Yu-Rou Tuan, Hao-Tang Tsui, Nicolas Ugrinovic, Kris Kitani, Xiaoxuan Ma
Involving before Evolving: A Vision for Trustworthy Enterprise Digital Twin Engineering
by Kérian Fiter, Adil Lagrou, Franck Dervault, Bentley Oakes
Beyond Word Error Rate: A Switch Aware Evaluation of ASR and Audio Language Models on English Yoruba Code-Switched Speech
by Chibuzor Okocha, Christan Earl Grant
Design Docs Are All You Need: An AI-native Machine-Learning Performance Tool
by Samuel Kushnir, Kimia Noorbakhsh, Kavya Sreedhar, Liqun Cheng, Ming Liu, Parthasarathy Ranganathan, Mohammad Alizadeh, Fred Kjolstad, Suvinay Subramanian