CyberMirror: Recovering Interpretable Attacker Reward Functions from Deep-RL Pentesting Agents via Inverse Reinforcement Learning

1. A demonstration-free pipeline on an emulation-friendly
range that transforms a PPO/MLP attacker into an
attacker-value detector without any human data, using
a sample-based Maximum-Entropy IRL reward recovery
algorithm with a separable objective and real partition
function, made to converge by (i) Markovian state-
occupancy features, (ii) horizon-matching the background
to the median demonstration length, and (iii) Ridge
regularization (λ = 0.3).

2. Empirical evidence that the recovered score differentiates
adversarial from benign states (AUC ≈ 1.0, 100% detect-
before-goal), and a detection-gated, shared-environment
duel in which this reward incentivizes an action-taking
defender, reported with its availability cost so that con-
tainment is not mixed up with free defense.

3. A transferable methodological finding on the trade-off be-
tween attack optimality and reward identifiability in IRL-
from-RL, explicitly stating the limitations and providing
a future roadmap.