Random-Key Optimizer with reinforcement learning for the Capacitated Multi-period Cutting Stock Problem with setup cost
| dc.contributor.author | Silva, Eduardo M. | |
| dc.contributor.author | Chaves, Antônio A. | |
| dc.contributor.author | de Araujo, Silvio A. [UNESP] | |
| dc.contributor.author | Jans, Raf | |
| dc.date.accessioned | 2026-06-18T18:59:02Z | |
| dc.date.issued | 2025-11-01 | |
| dc.description.abstract | This paper introduces a Random-Key Optimizer (RKO) procedure incorporating reinforcement learning to solve the One-Dimensional Multi-Period Cutting Stock Problem (MPCSP) with setup costs and capacity constraints. The MPCSP involves determining cutting plans for each period to meet customer demands, where inventory variables link consecutive periods. The RKO represents solutions as random-key vectors, which are decoded into feasible solutions for the MPCSP through a decoder process. During the optimization process, the RKO dynamically adapts its parameters using reinforcement learning. This framework integrates Biased Random-Key Genetic Algorithm (BRKGA), Particle Swarm Optimization (PSO), and Simulated Annealing (SA), all utilizing a unified decoder function. A novel penalization mechanism is also introduced within the decoder to handle infeasibilities effectively. The proposed RKO is evaluated on benchmark instances from the literature and compared against state-of-the-art methods, including a hybrid column generation heuristic and a dynamic programming-based heuristic. In addition, a new set of large-scale instances is introduced for further evaluation. Computational experiments reveal that the RKO employed by BRKGA consistently outperforms other solution methods in benchmark instances, delivering superior average solution quality. A sensitivity analysis is also conducted, examining the impact of setup costs and production capacity. Moreover, the study includes a comparative analysis of the RKO framework with and without reinforcement learning. | |
| dc.description.affiliation | Universidade Federal de São Paulo (UNIFESP-ICT), São José dos Campos, Brazil | |
| dc.description.affiliation | Universidade Estadual Paulista “Júlio de Mesquita Filho” (UNESP), São José do Rio Preto, Brazil | |
| dc.description.affiliation | GERAD Montréal, Canada | |
| dc.description.affiliation | HEC Montréal, Canada | |
| dc.description.affiliationUnesp | Universidade Estadual Paulista “Júlio de Mesquita Filho” (UNESP), São José do Rio Preto, Brazil | |
| dc.identifier | https://app.dimensions.ai/details/publication/pub.1189566122 | |
| dc.identifier.dimensions | pub.1189566122 | |
| dc.identifier.doi | 10.1016/j.cor.2025.107159 | |
| dc.identifier.issn | 0305-0548 | |
| dc.identifier.issn | 1873-765X | |
| dc.identifier.orcid | 0000-0003-1333-1426 | |
| dc.identifier.orcid | 0000-0001-5767-6798 | |
| dc.identifier.orcid | 0000-0002-4762-2048 | |
| dc.identifier.orcid | 0000-0001-8510-5677 | |
| dc.identifier.uri | https://hdl.handle.net/11449/326235 | |
| dc.publisher | Elsevier | |
| dc.relation.ispartof | Computers & Operations Research; v. 183; p. 107159 | |
| dc.rights.accessRights | Acesso restrito | pt |
| dc.rights.sourceRights | closed | |
| dc.source | Dimensions | |
| dc.title | Random-Key Optimizer with reinforcement learning for the Capacitated Multi-period Cutting Stock Problem with setup cost | |
| dc.type | Artigo | pt |
| dspace.entity.type | Publication | |
| relation.isOrgUnitOfPublication | 43c38943-bd6f-4fb6-a9a5-8482a1f632c0 | |
| relation.isOrgUnitOfPublication.latestForDiscovery | 43c38943-bd6f-4fb6-a9a5-8482a1f632c0 | |
| unesp.campus | Universidade Estadual Paulista (UNESP), Instituto de Biociências, Letras e Ciências Exatas, São José do Rio Preto | pt |

