Are the numbers right? Two questions about statistical software, one comparison engine: do independent implementations agree today, and did one of them used to be wrong in a way that reached published work.
Every stage is a subset of the one before. The drop is the finding, not an embarrassment — most changelog entries that sound like wrong numbers are not.
A regular expression proposed a category for all 163 exposed entries; each was then read and judged. Every correction is a measured disagreement, which is the validation a hand-coded gold set was going to provide.
One directory per bug, five pipeline stages, each leaving a durable artifact. Exposure is always a pair: scripts calling the function, and scripts also matching the bug's conditions probe.
Silent, result-changing, high severity, not yet verified — the queue the next verification pass draws from, ranked by corpus exposure.
Generated by kasauti dashboard. Every number here comes
from a command in the repository; none is hand-entered. Exposure counts are upper
bounds — a probe shows a script could have met the triggering
condition, never that it did.