Codeforces
AssessPlatforms
A competitive programming benchmark used as the problem set in the experiment.
Why it's here
Placed in Assess: 1 article(s) of evidence from 2 source(s), led by security coverage, with 1 in the last 30 days. Confidence 32%.
Evidence (1)
- 9Simon Willison·8/11/2026securityStudy Finds Hidden Reasoning Traces Leaking from Proprietary LLM APIs
Researchers claim they can reconstruct hidden reasoning traces from proprietary LLM APIs across models from OpenAI, Anthropic, and Google. The paper reports recovered privacy artifacts and secrets, including API keys, passwords, access tokens, and personal email addresses, from publicly available agent trajectories and reasoning blocks.