After multiple simulations comparing control performance, the value of β is set to −160 to ensure that the reward function exhibits an increasing trend during the iterative process.
← all excerpts
Long Short-Term Memory-Model Predictive Control Speed Prediction-Based Double Deep Q-Network Energy Management for Hybrid Electric Vehicle to Enhanced Fuel Economy.
1
—
—