RAPID (Real-time Arm Policy Distillation for Low-Latency Robot Control) is a research project for building a strong PPO robot-control Teacher, distilling it into a compact Student, and measuring whether quantization and structured pruning produce real CPU-latency gains without unacceptable control degradation.
- The original minimal PPO implementation is preserved as an immutable comparison baseline.
- The
Reacher-v5Minimal checkpoint was trained for 2.5 million environment steps, and the Improved checkpoint was trained for 3 million steps. - Gymnasium registers
-3.75as theReacher-v5reward threshold for considering the environment learned. - Over 100 deterministic episodes with seeds
0through99, the preserved Minimal checkpoint scores-4.0788 +/- 1.3084and does not reach the threshold. - Under the same protocol, the Improved checkpoint scores
-3.5086 +/- 1.2421, passes the Gymnasium threshold, and outperforms Minimal on 98 of 100 episode seeds. - Minimal and Improved PPO now share one teacher-training loop while keeping variant-specific transition storage separate.
- Minimal and Improved PPO also share one fixed-seed evaluation loop for directly comparable results.
- Pusher-v5 numbered experiments
000–014were screened under training seed 0, and Improved PPO versusexperiment_012was then evaluated across training seeds0–4. experiment_012improved the five-seed mean 5 cm success rate by+8.2percentage points, but missed the predefined+10-point gate. It was not promoted; Improved PPO remains the reference Teacher and 012 is retained only as a precision comparison condition.- Pusher evaluation now reports and saves return, mean/median/p90 final distance, and 5/8/10 cm success rates.
- Improved and numbered runs separate checkpoint training seeds from evaluation episode seeds and store independent replications below
seed_N. - Stored numbered-experiment settings can be replayed under a new training seed without changing the experiment number or depending on the current mutable config.
- PPO-CMA remains a separate ablation.
- Distillation, quantization, pruning, and common evaluation boundaries are scaffolded so they can be added without rewriting the baseline layout.
Minimal PPO baseline
|
Improved PPO reference Teacher -------- PPO-CMA ablation
|-- experiment_012 precision comparison (not promoted)
|
Teacher trajectory dataset
|
FP32 Student (BC / KD)
|
PTQ / QAT / structured pruning / combined variants
|
Common control, model-size, and CPU-latency evaluation
|
Aggregated tables and regenerated final figures
The required final comparison is:
- FP32 Teacher
- FP32 Student
- Student + PTQ or QAT
- Student + structured pruning
- Student + quantization + pruning
RAPID-Policy-Distillation/
|-- rapid/ # Importable research code
| |-- algorithms/ppo/
| | |-- minimal.py # Preserved minimal PPO baseline
| | `-- improved.py # GAE and minibatch Improved PPO
| |-- data/ # Teacher trajectory collection and datasets
| |-- distillation/ # Student models, BC/KD losses, and training
| |-- compression/
| | |-- quantization/ # PTQ, QAT, calibration, and integer export
| | `-- pruning/ # Structured pruning and recovery training
| |-- evaluation/ # Shared control, size, and latency metrics
| |-- envs/ # Environment factories and success criteria
| `-- utils/ # Seeds, metadata, checkpoint, and logging tools
|-- configs/ # Reusable component settings
| |-- ppo/ # PPO algorithm settings by variant
| |-- teacher/
| |-- data/
| |-- distillation/
| |-- compression/
| `-- evaluation/
|-- scripts/ # Thin command entry points by pipeline stage
| |-- teacher/
| | |-- train/ # Minimal, Improved, and experiment training commands
| | |-- eval/ # Matching evaluation commands
| | `-- utils/ # Shared adapters, config, environment, and run tracking
| |-- data/
| |-- student/
| |-- compression/
| |-- evaluation/
| `-- analysis/
|-- experiments/ # Versioned experiment manifests and hypotheses
|-- tests/
| |-- unit/
| `-- integration/
|-- artifacts/
| |-- baselines/ # Selected, versioned reproducibility assets
| |-- comparison_sets/ # Final checkpoints and evaluations for reported comparisons
| |-- experiment_records/ # Public logs and metadata for all Teacher experiments
| |-- runs/ # Generated checkpoints and run metadata
| |-- datasets/ # Generated trajectory datasets
| `-- exports/ # Generated deployable compressed models
|-- results/
| |-- tables/ # Final aggregate data
| |-- figures/ # Figures regenerated from final tables
| `-- summaries/ # Machine-readable comparison summaries
`-- docs/ # Architecture and experiment protocol
The names minimal_ppo, improved_ppo, and the planned ppo_cma are intentionally separate. New Teacher work must not silently change the baseline implementation, configuration, checkpoint, or logs.
Run all commands from the repository root.
The current local environment is Conda mujoco11 with Python 3.11.
conda activate mujoco11
python -m pip install -r requirements.txtConfirm that the active interpreter is the Conda environment before starting a long run:
python --version
python -c "import torch, gymnasium; print(torch.__version__, gymnasium.__version__)"| Variant | Algorithm config | Teacher config | Training command |
|---|---|---|---|
| Minimal PPO | configs/ppo/minimal.py |
configs/teacher/reacher_v5_minimal.py |
python -m scripts.teacher.train.minimal |
| Improved PPO | configs/ppo/improved.py |
configs/teacher/reacher_v5_improved.py |
python -m scripts.teacher.train.improved |
| Improved PPO | configs/ppo/improved.py |
configs/teacher/pusher_v5_improved.py |
python -m scripts.teacher.train.improved --config configs.teacher.pusher_v5_improved |
| Numbered ablation | configs/ppo/experiments.py |
configs/teacher/pusher_v5_experiments.py |
python -m scripts.teacher.train.experiments |
The PPO config controls rollout/update and optimization settings. The Teacher config controls the environment, total steps, logging and checkpoint frequencies, action standard-deviation schedule, seed, checkpoint run number, and normalized policy-action contract. Improved PPO accepts either a dotted config module or a repository-relative .py path through --config.
Configuration modules are loaded when training starts. Editing a config does not change a process that is already running; restart training to use new values.
# Immutable comparison implementation with a new run directory
python -m scripts.teacher.train.minimal
# GAE and minibatch implementation
python -m scripts.teacher.train.improved
# The same Improved PPO entry point, fully driven by the Pusher config
python -m scripts.teacher.train.improved `
--config configs.teacher.pusher_v5_improved
# Numbered Pusher ablation with a hypothesis recorded in metadata.json
python -m scripts.teacher.train.experiments `
--note "control run before changing one PPO parameter"
# Independent Improved Pusher replication
python -m scripts.teacher.train.improved `
--config configs.teacher.pusher_v5_improved `
--training-seed 1 `
--note "matched baseline replication"
# Replay the exact stored experiment 012 condition under training seed 1
python -m scripts.teacher.train.experiments `
--training-seed 1 `
--replicate-experiment 12 `
--source-seed 0 `
--note "multi-seed replication of experiment 012"Omitting --config preserves the original Reacher command. The shared environment builder reads the selected config in both training and evaluation; for Pusher it clips PPO actions in normalized [-1, 1] space and rescales them to the environment's [-2, 2] actuator range.
Both commands use the same environment loop, CSV logging, action-standard-deviation schedule, and checkpoint code. The adapters preserve the algorithm-specific transition contracts:
- Minimal PPO appends rewards and terminal flags to its legacy buffer.
- Improved PPO calls
store_step()before environment reset so truncated transitions can bootstrap from the final observation.
Generated runs are kept separate by PPO variant:
artifacts/runs/teacher/minimal_ppo/reacher_v5/seed_0/
|-- logs/
`-- checkpoints/
artifacts/runs/teacher/improved_ppo/reacher_v5/seed_0/
|-- logs/
`-- checkpoints/
artifacts/runs/teacher/improved_ppo/pusher_v5/seed_0/
|-- metadata.json
|-- logs/
|-- checkpoints/
`-- evaluations/
artifacts/runs/teacher/experiments_ppo/pusher_v5/seed_0/
`-- experiment_000/
|-- metadata.json
|-- logs/
|-- checkpoints/
`-- evaluations/
artifacts/runs/teacher/experiments_ppo/pusher_v5/seed_1/
`-- experiment_012/ # same condition, independent training seed
CSV logs contain episode,timestep,reward. The shared trainer prints average reward at print_freq, writes CSV data at log_freq, and saves model weights at save_model_freq.
Numbered experiments add the experiment number to the CSV filename and an experiment column to every row. Each metadata.json stores the experiment number, note, status, timestamps, effective PPO parameters, Teacher training/evaluation parameters, and artifact paths. Scheduled experiment checkpoints are retained by timestep as well as through one latest-checkpoint path.
The publication-oriented files are separated from the mutable run directories:
artifacts/comparison_sets/teacher/pusher_v5/improved_vs_experiment_012/
|-- manifest.json
|-- improved_baseline/ # reference Teacher, seeds 0-4
`-- experiment_012/ # precision comparison, seeds 0-4
artifacts/experiment_records/teacher/
|-- index.json
`-- ... # public CSV logs, evaluation JSON/CSV, and metadata
Only the reported Improved-versus-012 comparison set publishes model checkpoints. Other experiments publish their logs and metadata without model files.
To follow the newest Improved PPO CSV log in another PowerShell window:
$latestLog = Get-ChildItem artifacts\runs\teacher\improved_ppo\reacher_v5\seed_0\logs\*.csv |
Sort-Object LastWriteTime |
Select-Object -Last 1
Get-Content $latestLog.FullName -WaitTo plot the newest Improved PPO log and keep the graph updating during training:
python -m tests.plot_training_log --variant improved --watchFor a one-time image export without opening a graph window:
python -m tests.plot_training_log --variant improved `
--rolling-window 10 `
--save results\figures\improved-training.png `
--no-showPass a CSV path as the positional argument to plot a specific run instead of automatically selecting the newest log.
To compare several runs on the same timestep axis, edit the LOGS list near the top of tests/compare_training_logs.py, then run:
python -m tests.compare_training_logsEach LOGS entry is a (legend label, CSV path) pair. The default comparison uses the preserved Minimal PPO baseline and the current Improved PPO run. ROLLING_WINDOW, raw-log visibility, and an optional SAVE_PATH are configured beside the list.
Numbered Pusher experiments can be discovered and compared automatically from their experiment_### directories:
# Compare every numbered experiment
python -m scripts.analysis.compare_teacher_experiments
# Compare selected runs at the same training horizon
python -m scripts.analysis.compare_teacher_experiments `
--experiments 0 2 4 `
--max-timestep 1000000 `
--show-raw
# Export without opening a graph window
python -m scripts.analysis.compare_teacher_experiments `
--rolling-window 10 `
--save results\figures\pusher-experiments.png `
--no-showThe upper plot overlays smoothed learning curves on one timestep axis. The lower plot compares each run's last and best rolling reward. Labels use the recorded --note; when a note is absent, they show PPO parameter changes relative to the first selected experiment. Running and interrupted experiments are marked in the label and can still be plotted from their completed CSV rows.
For multi-seed work, use a new --training-seed rather than changing checkpoint_run. A replicated numbered condition keeps the original experiment number and is stored below a different seed_N; creating the same replicated experiment directory twice fails instead of silently allocating a different condition number. A standard non-replicated experiment still allocates the next number below the configured seed directory.
Stopping with Ctrl+C ends the process. Improved runs retain their latest scheduled checkpoint, while numbered experiments retain timestep checkpoints and a latest-checkpoint path. The trainer does not yet resume optimizer and rollout-buffer state.
Minimal and Improved PPO use the same evaluation loop and reporting protocol:
# Evaluate the preserved Minimal PPO baseline
python -m scripts.teacher.eval.minimal
# Evaluate the current Improved PPO run checkpoint
python -m scripts.teacher.eval.improved
# Evaluate the Pusher checkpoint with the same config used for training
python -m scripts.teacher.eval.improved `
--config configs.teacher.pusher_v5_improved `
--training-seed 0 `
--seed-start 100 `
--episodes 100 `
--save-metrics
# Evaluate the latest numbered experiment checkpoint
python -m scripts.teacher.eval.experiments
# Evaluate a saved checkpoint from experiment 3 at 500k timesteps
python -m scripts.teacher.eval.experiments `
--experiment 3 `
--timestep 500000
# Evaluate experiment 012 from training seed 1 on the common holdout
python -m scripts.teacher.eval.experiments `
--experiment 12 `
--training-seed 1 `
--seed-start 100 `
--episodes 100 `
--save-metrics
# Aggregate matched Improved and experiment 012 summaries across training seeds
python -m scripts.analysis.compare_teacher_seeds `
--experiment 12 `
--training-seeds 0 1 2 3 4 `
--seed-start 100 `
--episodes 100 `
--output-json results/summaries/pusher_exp012_5seed.json
# Run all unit and integration smoke tests
python -m unittest discover -vEvaluation defaults remain config-driven for legacy commands. For reported Pusher comparisons, explicitly select the checkpoint's training seed with --training-seed and use the shared deterministic holdout with --seed-start 100 --episodes 100 --save-metrics. The saved episode CSV and summary JSON are written below that run's evaluations/ directory. This avoids conflating checkpoint identity with episode initialization or comparing a deterministic checkpoint against stochastic training rewards.
The multi-seed comparison aggregates the summary mean from each independently trained checkpoint and reports the sample standard deviation across training seeds. It does not pool all evaluation episodes as if they were independent training runs. The completed Pusher comparison uses training seeds 0–4 and deterministic evaluation seeds 100–199 for every checkpoint.
| Metric | Improved PPO | Experiment 012 | 012 - Improved |
|---|---|---|---|
| Mean return | -26.1933 | -26.0376 | +0.1558 |
| Mean final distance | 8.5493 cm | 8.1357 cm | -0.4136 cm |
| Median final distance | 6.8361 cm | 6.1921 cm | -0.6440 cm |
| P90 final distance | 12.5186 cm | 12.7284 cm | +0.2098 cm |
| 5 cm success | 4.6% | 12.8% | +8.2%p |
| 8 cm success | 69.2% | 75.0% | +5.8%p |
| 10 cm success | 82.0% | 84.6% | +2.6%p |
Experiment 012 passed the other aggregate tolerances but missed the predefined 5 cm improvement gate by 1.8%p, with additional seed-dependent p90 variation. It is therefore not a new Improved version. Improved PPO remains the reference Teacher for downstream distillation, while 012 is used only to compare high-precision behavior and tail-distance trade-offs. The machine-readable aggregate is results/summaries/pusher_exp012_5seed.json.
Gymnasium registers -3.75 as the Reacher-v5 reward threshold. The project uses that value as the initial Teacher-pipeline promotion gate.
| Checkpoint | Training steps | Deterministic episodes | Mean return | Episode std | Gymnasium threshold |
|---|---|---|---|---|---|
| Minimal PPO baseline | 2,500,000 | 100 | -4.0788 | 1.3084 | Below |
| Improved PPO | 3,000,000 | 100 | -3.5086 | 1.2421 | Passed |
The 100-episode comparison uses the same seeds 0 through 99. Improved gains +0.5703 mean return over Minimal and performs better on 98 of the 100 paired episode seeds. This passes the Reacher Teacher-pipeline gate and supports proceeding to the Pusher Teacher experiment. It is evidence about these checkpoints; algorithm-level claims still require several independently trained seeds.
To reproduce the promotion-gate evaluation, set total_test_episodes=100 in both Reacher Teacher evaluation configs before running the two commands above. The default 30-episode setting remains useful for a faster checkpoint check.
The preserved Minimal PPO checkpoint and reference logs are under artifacts/baselines/teacher/minimal_ppo/reacher_v5/seed_0/. Training commands never write into this baseline directory.
- Evaluate all model variants on the same machine, environment version, episode seeds, success definition, and latency protocol.
- Report episodic return and success rate together with parameter count, serialized size, compression ratio, and CPU single-thread latency (
median,p95, andp99). - Split Teacher trajectory datasets by episode, not by individual transition.
- Keep fake-quantized accuracy separate from latency measured using a real exported integer runtime.
- Generate final plots from stored aggregate tables. Temporary GIFs, videos, and exploratory plots do not belong in source control.
See docs/architecture.md and docs/experiment_protocol.md for the detailed boundaries and promotion criteria.
- Proximal Policy Optimization Algorithms
- What Matters In On-Policy Reinforcement Learning?
- PPO-CMA: Proximal Policy Optimization with Covariance Matrix Adaptation
This project is distributed under the MIT License.
The minimal PPO baseline is derived from PPO-PyTorch by Nikhil Barhate. The original copyright notice is retained in LICENSE. RAPID project modifications are copyright 2026 Dongha Kim (김동하) and 윤기찬.