ARTFEED — Contemporary Art Intelligence

Differentiable Projection Enforces Hard Constraints in Deep RL for Inventory Policies

other · 2026-08-04

A new arXiv preprint (2608.02343) introduces a method for learning feasible inventory policies in constrained sequential decision processes. The approach embeds a differentiable convex optimization module within a deep reinforcement learning (DRL) policy. A neural network proposes continuous action targets, which are then projected onto a relaxed feasible set via a quadratic program. A dual-informed integer mapping restores integrality while preserving feasibility. The policy is trained end-to-end using pathwise gradients from sampled trajectories in a differentiable simulator. This addresses limitations of existing DRL methods that either penalize constraints or rely on feasibility mechanisms that fail when constraints interact. The work targets operational problems with large, combinatorial action spaces and interdependent feasibility constraints, such as those modeled by mixed-integer linear programs (MILPs), which scale poorly in stochastic environments. The paper is announced as a new submission on arXiv.

Key facts

  • arXiv:2608.02343 is a new preprint.
  • The method uses a differentiable convex optimization module inside a DRL policy.
  • A neural network proposes continuous action targets.
  • A quadratic program projects actions onto the relaxed feasible set.
  • A dual-informed integer mapping restores integrality.
  • Training uses pathwise gradients from sampled trajectories.
  • The approach handles hard constraints without penalty methods.
  • It targets operational problems with large action spaces and interdependent constraints.

Entities

Institutions

  • arXiv

Sources