-
Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs
Paper • 2605.09063 • Published • 82 -
NanoResearch: Co-Evolving Skills, Memory, and Policy for Personalized Research Automation
Paper • 2605.10813 • Published • 15 -
AutoMedBench: Towards Medical AutoResearch with Agentic AI Models
Paper • 2606.01961 • Published • 26
Dingjie Song
songdj
AI & ML interests
None yet
Recent Activity
upvoted a paper about 4 hours ago
Opera: A Verbal Critic Framework for Long-horizon Coding Agents updated a model 4 days ago
songdj/m-638 published a dataset 4 days ago
songdj/a-426Organizations
None yet