News
AI, cybersecurity, machine learning, scraping & programming news.
33 results on this page · clear filters
Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence
When an LLM inside an agent workflow produces a recommendation or judgment, it often comes with an explanation naming the factors that drove the decision. Operators use these explanations to monitor systems, diagnose errors, or de...
Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models
When an LLM reports a molecular property with a median absolute error of 0.025 kcal/mol on the FreeSolv benchmark, that number is far below the 0.6 kcal/mol experimental uncertainty assigned to the measurements. No model can predi...
When LLM Decompilers Recompile More and Preserve Less
When LLM Decompilers Pass Every Test Yet Rewrite Your Code Decompilation turns compiled binaries back into readable source code, and the stakes are high: security analysts depend on it to find vulnerabilities in malware, reverse-e...