Portrait of Zhenhao Zhang

Zhenhao Zhang

张臻昊
M.S. Student in Artificial Intelligence
Columbia University

About

I am an M.S. student in Artificial Intelligence at Columbia University (concentration: AI and Advanced Computing). I received my B.S. in Computer Science and B.S. in Business (Information Systems) from the University of Rochester in 2025, graduating cum laude with Highest Distinction in Computer Science. Since 2025, I have been working with Dr. Linchao Zhu at the College of Computer Science and Technology, Zhejiang University — first on agent memory, and since 2026 on computer-use agents.

My research focuses on AI agents, multimodal learning, and reinforcement learning. I have worked on computer-use agents, agent memory, reward modeling, reasoning distillation, supervised fine-tuning (SFT), and reinforcement learning for agent training.

I’m particularly interested in building generalizable agents that can reason, interact with complex environments, learn from feedback, and improve through experience.

Research

CaptchaArena: A Large-Scale, Fine-Grained Dataset for Training Computer-Use Agents on Interactive CAPTCHAs
Zhenhao Zhang*, Zhaoyu Fan, Haohan Ying, Jingwen Hu, Hancen Fan, Junhao Zhou, Zitian Chen, Linchao Zhu†  (*Project lead, †Corresponding author)
arXiv 2026 · Under review
Project Lead · Zhejiang University, College of Computer Science and Technology · Advisor: Dr. Linchao Zhu · Jan 2026 – Present
Contributions
  • Led an 8-person team to build CaptchaArena, the first large-scale, fine-grained training dataset for interactive CAPTCHA solving: 50K puzzles across 20 types and 5 interaction modes (single-click, multi-click, arrow-cycle, real-time, text-entry), each with an executable reference solution replayed in a real browser and accepted by the environment verifier, plus pixel-mask grading for irregular targets.
  • Built a reasoning-annotation pipeline that turns 50K verified screenshot–action trajectories into 46K step-by-step reasoning-annotated trajectories, using a teacher model, a cross-family VLM judge (consistency, no hindsight, decisiveness), feedback-driven regeneration, and stronger-model verification.
  • Trained CaptchaAgent, a single Qwen3.5-9B policy for all 20 types: supervised fine-tuning on 37.6K per-turn samples raises Pass@1 from 11.4 to 70.5, and GRPO with the environment verifier as reward (no reward model or human labels) raises it to 71.7 — above the strongest open-weight GUI agent (35.2) and closed-source model (69.2); humans reach 94.1.
  • RL also improves transfer to external benchmarks (Open CaptchaWorld 47.2 → 51.0, Halligan 13.6 → 20.0); conducted per-type error analysis showing remaining failures concentrate in non-submission, weak grounding, and unstable execution.
  • Engineered the multimodal SFT and GRPO training infrastructure with DeepSpeed ZeRO-3, Liger FLCE, distributed checkpointing, and live-environment rewards for long-context training.
SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents
Yang Wan, Zhenhao Zhang, Jierui Wang, Linchao Zhu
arXiv 2026
Research Assistant · Zhejiang University, College of Computer Science and Technology · Advisor: Dr. Linchao Zhu · Jan 2026 – Aug 2026
Contributions
  • Built and manually annotated a 300+ trajectory cross-platform benchmark for long-horizon computer-use agents, providing trajectory-level verdicts and dense step-level supervision across web, desktop, and mobile environments.
  • Designed fine-grained reward annotation criteria and quality-control workflows covering milestones, correct decisions, error correction, off-path behavior, and harmful actions, and analyzed model–human agreement at the step level.
  • Conducted human evaluation and systematic error analysis of rule-based and model-based reward signals, identifying rule–human disagreement across multiple CUA environments and motivating learned reward models for long-horizon agent evaluation.
  • Contributed to the distillation and SFT training of SeekJudge-9B, analyzing failure modes in reasoning behavior, output formatting, and termination control, and iterating on training data and supervision strategies for the unified reward-model backbone.
Mitigating Conversational Inertia in Multi-Turn Agents
Yang Wan, Zheng Cao, Zhenhao Zhang, Zhengwen Zeng, Shuheng Shen, Changhua Meng, Linchao Zhu
ICML 2026
Research Assistant · Zhejiang University, College of Computer Science and Technology · Advisor: Dr. Linchao Zhu · Feb 2025 – Aug 2026
Contributions
  • Implemented task-specific trajectory summarization for long-horizon LLM agents across Maze, WebShop, Wordle, SciWorld, and TextCraft, compressing multi-turn observation–action histories into compact memory while preserving task-critical state and progress.
  • Designed and optimized environment-aware memory representations and summarization prompts, selectively retaining navigation states, completed subgoals, constraints, previous actions, and unresolved information for downstream decision-making.
  • Benchmarked summarization against alternative context-management strategies across heterogeneous AgentGym environments, analyzing the trade-off between history retention and conversational inertia; training-free summary prompts improved Maze success from 85.5% to 90.0%.
  • Conducted failure and behavioral analysis of long-context agents, building attention diagnostics that link system-token ratio → system attention → task success, and studying action repetition, exploration, and error propagation.
Political Symbols and Urban Amenities: The Spatial Logic of Flag Placement in Istanbul
Cantay Çalışkan, Yi Ren, Yiheng Yao, Zhenhao Zhang, Qike Jiang, Yuewen Yan
SSRN 2026
Undergraduate Research Assistant · University of Rochester, Goergen Institute for Data Science and Artificial Intelligence (DSCC 279 Independent Research) · Advisor: Prof. Cantay Çalışkan · Aug 2025 – Jan 2026
Contributions
  • Built a large-scale geospatial data pipeline over ~748K Google Street View panoramas — image cleaning, coordinate alignment, spatial matching, and regional aggregation — to construct an analysis-ready urban visual dataset.
  • Designed a multi-scale spatial feature extraction framework using nested geographic buffers and distance-weighted aggregation to transform panorama-level visual signals into constituency-level statistical representations.
  • Developed spatial distribution metrics for political-symbol visibility, applying Earth Mover's Distance (EMD) to quantify structural differences in visual-signal distributions across urban districts.
  • Conducted cross-region spatial analysis linking visual signals with urban geographic structure, translating large-scale image observations into interpretable district-level features for downstream statistical analysis.

Education

Columbia UniversityNew York, NY
M.S. in Artificial Intelligence · Concentration: AI and Advanced ComputingAug 2026 – Dec 2028 (expected)
University of RochesterRochester, NY
GPA 3.85/4.00 · Dean's List (2023–2025)
B.S. in Computer Science — Cum Laude, Highest Distinction in Computer ScienceAug 2022 – Dec 2025 (graduated early)
B.S. in Business: Information Systems — Cum Laude

Awards

GIDS Biomedical Data Science Hackathon — Prize Winner · 2nd PlaceAug 2026
GIDS Biomedical Data Science Hackathon — Prize WinnerAug 2025
GIDS Biomedical Data Science Hackathon — Prize WinnerAug 2024
Finnov8 Hackathon — 2nd Place OverallMar 2024

Teaching

Undergraduate Teaching Assistant, CSC 280: Computer Models and Limitations — Department of Computer Science, University of Rochester, Fall 2025