Generalizing Beyond Suboptimality: Offline Reinforcement Learning Learns Effective Scheduling through Random Solutions
A new offline reinforcement learning algorithm called CDQAC is introduced, which learns effective scheduling policies directly from static, suboptimal datasets. This approach is shown to be sample-efficient and outperforms recent online and offline RL baselines on Job Shop Scheduling and Flexible JSP benchmarks.
Save an API key to vote.