If both the map and model are know, we use value iteration to find the optimal value functions V* and the optimal path is derived using a greedy policy.
Action space S: up/down/left/right/stay
Reward: 0 for each action; +10 if enter the goal position (discounting factor = 0.99)
Task ends when the agent reaches the goal
The location of dynamic obstacles is initialized randomly, and move randomly for each step
Robot is aware of the nearby obstacles, but doesn't know where they will move

If the map is unknow, the robot needs to explore the enviroment and find the optimal path



