Skip to content
View yuchenwang3's full-sized avatar
:octocat:
Focusing
:octocat:
Focusing

Block or report yuchenwang3

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
yuchenwang3/README.md

Yuchen (Ean) Wang — handwritten typing signature

Agentic post-training & ML systems.
M.S. CS @illinois · Research intern @Accio-Lab / @alibaba.

Website CV Scholar LinkedIn Email WeChat · eangyc

Research

Occamy-1.0: official logo

35B-A3B agent model for long-horizon tool use.
Execution-grounded data · Agentic post-training

Report Demo Model Code

CineFlow: figure from the paper

Dependency-driven parallel video generation.
1.7–5.5× end-to-end speedup in the reported evaluation

Project Paper

Dynamic Prefill: figure from the project report

Adaptive batching and prompt packing for LLM serving.
Up to 20% lower TTFT on the reported traces

Report Code

CUDA Attention RL for Legal Reasoning

Open-source contributions

9 merged · 1 adopted solution · 20 open · 13 projects

Highlights

vllm-project organization avatar
vLLM

#54699 · merged Remove full-weight copies during MoE loading; conversion peak 7.88 → 3.94 GiB in the exact-shape TP2 benchmark

NVIDIA-NeMo organization avatar
NeMo RL

#3943 · open Bypass driver tensor materialization in distillation; 4.4–5.3× faster transfers in a controlled Ray benchmark

NVIDIA organization avatar
Megatron-LM

#5396 · open Fuse GDN Q/K normalization to remove an extra backward activation buffer

#5463 · open Enable selective Mamba mixer recompute to save activation memory without full-layer recomputation

Dao-AILab organization avatar
FlashAttention

#2507 · merged Prevent redundant backward-kernel recompilation with stable cache keys, without CPU–GPU sync
Solution adopted by the PR author

flashinfer-ai organization avatar
FlashInfer

#4984 · merged Restore K/V calibration in FP8 KV prefill, correcting silently mis-scaled attention outputs

sgl-project organization avatar
SGLang

#39765 · open Publish Mamba cache updates before dependent batches capture stale KV mappings

NousResearch organization avatar
Hermes Agent

#100693 · open Resolve local schema references so nested tool arguments reach handlers as objects, not JSON strings

More contributions

ms-swift · 4 contributions
modelscope organization avatar
ms-swift

#9598 · merged Add order-preserving packing

#9602 · merged Warm up NCCL before training

#9599 · merged Pass through Muon Nesterov settings

#9591 · merged Expose Muon coefficient selection

vime · 1 contribution
vllm-project organization avatar
vime

#337 · merged Forward recompute flags; fix hybrid models

Emerging Optimizers · 1 contribution
NVIDIA-NeMo organization avatar
Emerging Optimizers

#230 · merged Keep Muon scale-invariant

NeMo Gym · 1 contribution
NVIDIA-NeMo organization avatar
NeMo Gym

#2726 · merged Preserve HTTP errors across process boundaries
Co-author

Megatron-LM · 3 contributions
NVIDIA organization avatar
Megatron-LM

#5400 · open Route GDN input projections to Adam

#5431 · open Exclude GDN input projections from global clipping

#5395 · open Skip gradient clipping for Muon

NeMo RL · 3 contributions
NVIDIA-NeMo organization avatar
NeMo RL

#4193 · open Render evaluation prompts as complete conversations

#4176 · open Unify worker selection through configuration

#2962 · open Sanitize non-finite async log probabilities

SGLang · 3 contributions
sgl-project organization avatar
SGLang

#40103 · open Reject developer messages silently dropped by templates

#38063 · open Explain cold MXFP4 JIT startup

#31621 · open Honor weight-check exclusions during reset

verl · 2 contributions
verl-project organization avatar
verl

#7906 · open Track response truncation across context limits

#7597 · open Validate actor FSDP strategy

Hermes Agent · 3 contributions
NousResearch organization avatar
Hermes Agent

#113511 · open Control partial-stream continuation for batch evaluation

#113538 · open Clarify API retry budgets and streaming defaults

#102549 · open Make SSH reconnects race-safe

TRL · 1 contribution
huggingface organization avatar
TRL

#7294 · open Fix async checkpoint resume after stale rollout drops

All contributions ↗ Engineering notes ↗

On GitHub

Follow on GitHub Stars on my repositories Explore my open pull requests Profile views

GitHub contribution rhythm over 26 weeks Primary languages of my public non-fork repositories

Contribution trail

Snake animation of my GitHub contribution history

Pinned Loading

  1. Accio-Lab/occamy Accio-Lab/occamy Public

    Occamy model repo

    JavaScript 106 7

  2. vllm-project/vllm vllm-project/vllm Public

    A high-throughput and memory-efficient inference and serving engine for LLMs

    Python 92.7k 22.7k

  3. sgl-project/sglang sgl-project/sglang Public

    SGLang is a high-performance serving framework for large language models and multimodal models.

    Python 36.4k 9.2k

  4. NVIDIA/Megatron-LM NVIDIA/Megatron-LM Public

    Ongoing research training transformer models at scale

    Python 18k 4.6k

  5. NousResearch/hermes-agent NousResearch/hermes-agent Public

    The agent that grows with you

    Python 249k 52.9k

  6. verl-project/verl verl-project/verl Public

    verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

    Python 23.6k 4.6k