Repository navigation
Question: direct RL with larger RGB inputs / frozen visual encoders in mjlab (without distillation) #936
Unanswered
weristeddy
asked this question in
Q&A
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Hi, I had a question about the vision setup in mjlab.
I am using the YAM lift-cube RGB task for a master’s thesis. The default small end-to-end CNN works, but for sim-to-real reasons I am interested in using larger RGB inputs, ideally at least 128×128.
From the mjlab paper, my understanding is that high-fidelity RGB rendering is mostly out of scope, and that one common workflow is to train privileged policies first and then distill to vision-based controllers using external rendering. In my case, though, I would specifically like to avoid distillation if possible.
So I wanted to ask two related things.
1. Larger RGB inputs in general
I first wanted to ask about image resolution itself: is training with larger RGB inputs such as 128×128 something mjlab is meant to support in practice?
My main motivation here is eventual real-world deployment, so I am interested in an observation setup that is a bit closer to what I would want for sim-to-real experiments.
2. Pretrained visual encoders
Separately, when I replace the default small CNN with a larger frozen pretrained encoder, training becomes extremely slow because visual encoding dominates the cost, especially with high parallelism.
Have you run into similar issues yourselves, or tried similar setups internally? If yes, I would be very interested in what your experience was and whether you found any practical way to make this work.
More generally, if distillation is not desired, is there any recommended approach for either of these settings?
For example:
I am mainly trying to understand whether this kind of setup is:
Any pointers would be very helpful.
All reactions