Figure 6 shows that the average episode reward shows an increasing trend, verifying the agents indeed learned to generate better solutions.
← all excerpts
Intelligent Decision-Making of Scheduling for Dynamic Permutation Flowshop via Deep Reinforcement Learning.
1
—
—