목적
RL 환경이 실제 crew pairing 문제 정의를 반영하도록 state와 constraint를 확장하고,
reward를 재설계하여 curriculum learning 기반 학습 가능성을 검증
작업 내용
- state representation을 실제 scheduling 요소 기반으로 확장
(time, location, duty accumulation 등)
- constraint를 명시적 feature 및 masking 구조로 재정의
- reward function을 단계별 학습이 가능하도록 재설계
- curriculum learning 전략 적용 (difficulty / constraint 점진적 증가)
실험 목표
- 단순 설정 대비 학습 안정성 비교
- reward shaping이 policy 학습에 미치는 영향 분석
- curriculum 적용 시 수렴 여부 확인
목적
RL 환경이 실제 crew pairing 문제 정의를 반영하도록 state와 constraint를 확장하고,
reward를 재설계하여 curriculum learning 기반 학습 가능성을 검증
작업 내용
(time, location, duty accumulation 등)
실험 목표