Learning Equitable and Efficient Shift Rosters with Multi-Objective Deep Reinforcement Learning and Counterfactual Fairness Constraints in Operations
Published 2026-08-13
Keywords
- Adverse impact,
- Counterfactual fairness,
- Deep reinforcement learning,
- Multi-objective optimization,
- Workforce rostering
Copyright (c) 2026 International Journal of Industrial Engineering and Management

This work is licensed under a Creative Commons Attribution 4.0 International License.
Abstract
Shift rostering systems often optimize cost and coverage while treating fairness as a secondary objective. This study develops a multi-objective deep reinforcement learning framework for workforce scheduling that integrates Proximal Policy Optimization with a graph neural network encoder, a Conditional Value-at-Risk term for limiting extreme burden concentration, and a counterfactual fairness regularizer for reducing demographic sensitivity in assignment decisions. The framework was evaluated on 18 months of workforce data from a tertiary hospital and an industrial facility in Uzbekistan. Relative to mixed-integer linear programming and vanilla reinforcement learning baselines, the proposed model reduced the burden Gini coefficient by 41.3%, narrowed weekend disparity by 36.5%, and achieved adverse impact ratios above 0.97 for gender and age. These equity gains were obtained with a 1.8% increase in operational cost and a 9.4% reduction in understaffing minutes under demand shocks. The results indicate that fairness constraints can be embedded directly into policy learning while preserving service levels and improving resilience in high-constraint scheduling environments.
Article history: Received (December 3, 2025); Revised (April 17, 2026); Accepted (May 13, 2026); Published online (August 13, 2026)