Junjie Xiong

Projects

rLLM

RL for Language Models

Reinforcement learning infrastructure and post-training systems

I am interested in the practical stack behind RL for LLMs: rollout generation, reward signals, policy optimization, evaluation, and debugging. I have been exploring open-source RL-for-LLM tooling including rLLM, with an eye toward making post-training more understandable and reproducible.

Agents' Last Exam (ALE)

App & Design Team contributor · In development

I contribute to the design and development of realistic, reproducible computer-use tasks for ALE, helping evaluate agents on complex workflows across different environments.

Towards Expert Financial QA via Self-Improving RAG

Accepted at AFA Workshop @ ICLR 2026

A judge-driven multi-agent retrieval system for financial document QA. The system decomposes answering into retrieval, reasoning, and judging agents, then uses feedback-driven retries to recover from weak answers. I see this as a bridge between retrieval systems and RL-style improvement loops: feedback, credit assignment, and iterative correction.

Vulci3000

Archaeological fieldwork and digital reconstruction

My contribution: Contributed to archaeological fieldwork using ground-penetrating radar, LiDAR, 3D scanning, infrared imaging, spatial data analysis, and Gaussian splatting to document and digitally reconstruct the site.

Broader project methods: GPR, magnetometry and electrical resistivity; drone LiDAR, multispectral and thermal imaging; photogrammetry and point clouds; GIS, georeferenced mapping, elevation models and vegetation indices; VR and digital archives.