An open-source Python tool from Pathway that generates ARC-AGI-1-style reasoning tasks matching the distribution of public evaluation sets. It creates private, unseen tasks in a standard format for comparing model performance beyond public benchmark familiarity.