Skip to content

feat: Add RLlib BC and MARWIL - #38

Open
noamonti748 wants to merge 10 commits into
mainfrom
noamonti748/rllib-imitation-learning
Open

noamonti748 wants to merge 10 commits into
mainfrom
noamonti748/rllib-imitation-learning

Conversation

@noamonti748

@noamonti748 noamonti748 commented Aug 21, 2026 •

Copy link
Copy Markdown

Summary

The PR enables the use of RLlib's BC and MARWIL with Schola.

  • Can connect to a running Imitation Environment, automatically collect data to parquet episodes, and perform training.

Related issues

None.

Type of change

  • New feature or enhancement
  • Documentation

Changes

  • Add BC and MARWIL through schola rllib bc and schola rllib marwil
  • Store demonstrations as RLlib Parquet episodes
  • Add a collector that records gameplay from Unreal though the existing simulator subcommands
  • Run collection and training in one command. Can also train from an existing dataset with --input
  • Extend the RLlib train CLI, checkpouint export, and command template for offline workflows
  • Add an [offline] pip extra for RLlib offline dependencies
  • Add an imitation learning guide and tests for offline data, CLI, and training

Testing

Python (Resources/python, Test)

  • Installed test dependencies: pip install --group test -e "./Resources/python[all]"
  • Ran: python -m pytest Test --import-mode=importlib -n 0

Other Tests

  • Tested end-to-end pipeline on a version of the Basic Environment adapted for IL.
  • Tested on a Python process running a cartpole example. One over gRPC, and one that writes to a Parquet set directly.

Checklist

  • Code follows project style (Unreal coding standard for C++; Black for Python)
  • Comments / docstrings added or updated where behavior is non-obvious
  • README or Sphinx docs updated if user-facing behavior changed
  • No unrelated changes included in this PR
  • Appropriate license headers added to new files

@noamonti748 noamonti748 changed the title Noamonti748/rllib imitation learning [feat] Add RLlib BC and MARWIL Aug 21, 2026
@noamonti748 noamonti748 self-assigned this Aug 21, 2026
@noamonti748 noamonti748 added the enhancement New feature or request label Aug 21, 2026
@noamonti748 noamonti748 changed the title [feat] Add RLlib BC and MARWIL feat: Add RLlib BC and MARWIL Aug 21, 2026
@noamonti748
noamonti748 marked this pull request as ready for review August 21, 2026 16:14
Comment thread Docs/Sphinx/guides/imitation_learning.rst Outdated
Comment thread Docs/Sphinx/guides/imitation_learning.rst Outdated
Comment thread Docs/Sphinx/guides/imitation_learning.rst Outdated

"""
Cyclopts template for generating Schola subcommands (multi-simulator and algorithm dispatch).
"""

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why are you modifying command_template.py in this PR? This code has proven difficult for AI to interpret, so there is a good chance these changes were unnecessary, or extremely complicated for what should be a small change.

What issue does the changes to command_template solve?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It's for sure overkill and will cut back, but the one thing I think it solves is that it allows us to run without binding to a simulator, because the default template binds to external by default whenever a simulator command is not included. Would it be ok to just isolate this part and kill the rest (should be < 50 lines I think)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The command template should support this without changes (0-N algorithms, 0-N simulators). This case has tests as well so should be working. If you want to change the default sim, just override the function in the child class (it should be ignored when there is no supplied simulators).

If you really want to change it, making it str or None to indicate No Default should be fine. In general, the actual structure is rather delicate and bolting on features rather than building them in properly is error prone.

Comment thread Resources/python/schola/scripts/rllib/offline_train.py Outdated
@@ -1,29 +1,43 @@
# Copyright (c) 2024-2025 Advanced Micro Devices, Inc. All Rights Reserved.
# Copyright (c) 2024-2026 Advanced Micro Devices, Inc. All Rights Reserved.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a bit of a mess, seems like a separate offline-train command is the way to go, alongside a collect command with functionality matching minari.

@noamonti748 noamonti748 Aug 21, 2026 •

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Split commands into these.

)


def load_rl_module_from_algorithm_checkpoint(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does another version of this code exist elsewhere already? I believe this functionality is required by the export command (to load the policy before converting it to a Schola Model API).

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep, that's true, but I think the issue is that RLModule load, where the rllib export goes through load_rl_module_from_algorithm_checkpoint or MultiRLModule.from_checkpoint is not shared between them. And of course, the load_rl... functionality is different.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What about for eval? The comment here says it shares layout handling with eval.

It may also be possible to make eval depend on this, ideally, we keep the number of distinct handlers for loading RLLib objects down lol. Between eval and rllib we already have a lot.

Feel free to make more drastic changes here if you see a way to make everything fit nicely. (e.g a generic Load function that everything else relies on in different places)

Comment thread Docs/Sphinx/guides/imitation_learning.rst Outdated
Comment thread Resources/python/schola/rllib/export.py Outdated
@@ -0,0 +1,480 @@
# Copyright (c) 2026 Advanced Micro Devices, Inc. All Rights Reserved.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How much of this can be outsourced to RLLIB?
As an example, see the JSON Writer for writing the dataset to a file.

JSON Writer
Instructions Recommending Ray Data tools

@noamonti748 noamonti748 Aug 21, 2026 •

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

JSONWriter/DatasetWriter seems to use the old-stack data format.
We can use the Parquet + msgpack over write_parquet though.
Also the SingleAgentEpisode

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good to me.

My philosophy here is that it will be easier in the long run to use the rllib tools, rather than reimplement, where possible since any re-implementations will be coupled to the output of the functions. If the output format changes, we will need to adjust our tools to match.

Noah Monti added 3 commits August 21, 2026 17:01
- Update imitation_learning guide for generic BC/IL with Minari
- `command_template` kept to original standard
- Shared checkpoint loading path
- Use RLlib functions for dataset writing
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants