You're measuring with a micrometer screw gauge and then drawing the line with a fat marker pen. A combat box was typically 500ft deep x 600 ft in altitude x 2500 ft in width. All of those dimensions impact fall of bombs onto the target area. How can CEP be solely about the aiming device when only one aircraft in such a large area was actually aiming at the target? Poor formation keeping will increase the size of the combat box and hence reduce bombing precision and accuracy, and that has nothing to do with the bomb sight.
Evaluating the accuracy of the bomb sight is done during testing which, in the case of the Norden system, led to the "bomb in a pickle barrel" claims (results which couldn't be replicated on operations for a host of reasons). However, the USSBS isn't about evaluating the effectiveness of the bomb sight. It's about assessing the effectiveness of the bombing campaign, which is a broader capability issue rather than a narrow focus on equipment. In modern US parlance, we'd talk about DOTmLPF-P. It's clear that US doctrine for daylight precision bombing changed from 1941 thru 1944, as did training, processes and everything else. You can have the best bomb sight in the world but if your crews can't get to the target reliably then, from a capability perspective, you're not successful.
I still say that discounting the results for formations which achieved less than 5% of bombs within 1000ft of the aim point is skewing the data, particularly given the size of the combat box.