counterfactual invariance

Identity or Prompt Noise? A Calibrated Invariance Audit of LLM Code Generation

Identity cues are irrelevant to a fixed programming specification, but raw counterfactual differences can arise from unequal samples and prompt wording. We audit 30.73 million executed Python generations from seven checkpoints on HumanEval+ and …