Publications
These papers gather around three questions: how signals can be described, how space and time can be programmed, and how agents behave when tasks become open-ended.
Conference Papers
2026
ICLR 2026
An instruction-tuned vision-language model and benchmark for detailed, generative clinical EEG interpretation.
Preprints and Benchmarks
2026
arXiv preprint
A large-scale benchmark for testing whether language models can generate runnable code that also produces spatially coherent animations.
2026
arXiv preprint
A bilingual benchmark of complex long-horizon academic tasks sourced from university students’ real workflows.