00 / 个人介绍 · About Me
我一直从事 Linux 系统与内核驱动开发、高并发通信架构设计、区块链密码学算法的 HPC 算子开发,以及机器学习推理系统的开发与性能优化。此外,我也曾从事自动驾驶通信架构开发、规划与控制算法开发,以及相关系统的性能优化。
业余时间,我也持续探索 LLM 推理、训练与强化学习(RL)相关技术。
My work spans Linux systems and kernel driver development, highly concurrent communication architectures, HPC kernels for blockchain cryptographic algorithms, and the development and performance optimization of machine learning inference systems. Previously, I developed communication architectures and planning and control algorithms for autonomous driving, and optimized the performance of these systems.
In my spare time, I also explore LLM inference, training, and reinforcement learning (RL).
技术栈 / Tech stack:C++ · C · Rust · Python · CUDA · Triton
熟悉的框架 / Familiar frameworks:vLLM · PyTorch
目前正在学习 RL 后训练的相关算法与框架,包括 vime、miles,以及训练框架 Megatron;同时深入探索集群通信、RDMA 通信,以及 Triton、LLVM、MLIR 等编译器相关技术。
I am currently studying algorithms and frameworks for RL-based post-training, including vime and miles, as well as the Megatron training framework. I am also exploring cluster communication, RDMA, and compiler technologies such as Triton, LLVM, and MLIR.
未来愿景 / Future vision
未来,我将持续深耕系统性能优化领域,致力于打造追求极致性能的基础设施,覆盖 LLM、RL、自动驾驶与传统系统,并努力成长为系统性能优化领域的专家。
I will continue to focus on systems performance optimization, building infrastructure that pushes performance to its limits across LLMs, reinforcement learning, autonomous driving, and traditional systems. My long-term goal is to become an expert in systems performance optimization.
04 / Open Source Contributions
Public contribution events in the latest 100-event snapshot. Forks and stars are excluded.
| Project | Contribution |
|---|---|
| RL-Align/RL-Kernel | MAINTAINER |
| vllm-project/vllm | PUBLIC CONTRIBUTIONS |
Featured Repositories
ncu-cuda-profiling-skillShell★ 126STARS
把 CUDA 性能分析融入 AI 编码工作流:让 Agent 借助 Nsight Compute 采集真实指标,辅助定位 kernel 瓶颈并整理优化建议。
Bring CUDA profiling into AI coding workflows: help agents collect real Nsight Compute metrics, investigate kernel bottlenecks, and organize optimization recommendations.
采集到报告:脚本串联 NCU 采集、CSV 导出、逐 kernel 摘要和 Markdown 分析报告。
From profiling to reports: automate NCU collection, CSV export, per-kernel summaries, and Markdown analysis.
基于指标的初筛:结合 DRAM、L1/TEX、SM 吞吐率与 Occupancy 指标给出规则诊断、判断依据及优化建议。
Metric-based triage: use DRAM, L1/TEX, SM throughput, and occupancy metrics to produce rule-based diagnoses, reasoning, and optimization suggestions.
适配多种 Agent:提供 Kimi、Claude Code、Cursor、Codex 安装入口,并支持环境检查和已有 NCU 报告分析。
Multiple agent integrations: installation options for Kimi, Claude Code, Cursor, and Codex, plus environment checks and analysis of existing NCU reports.
05 / Recent Activity
| Date (UTC) | Activity | Project |
|---|---|---|
| 2026-09-12 | pushed code | maxiaosong1124/RL-Kernel |
| 2026-09-12 | pushed code | maxiaosong1124/RL-Kernel |
| 2026-09-14 | pushed code | maxiaosong1124/RL-Kernel |
| 2026-09-12 | pushed code | maxiaosong1124/RL-Kernel |
| 2026-09-11 | pushed code | maxiaosong1124/maxiaosong1124 |
| 2026-09-11 | pushed code | maxiaosong1124/maxiaosong1124 |
| 2026-09-12 | created branch | maxiaosong1124/RL-Kernel |
| 2026-09-12 | starred repository | RL-Align/rl-align-website |
Public data snapshot: 2026-09-15T05:07:59+00:00. Activity is limited to the latest 100 public events; it is not a complete contribution history.