AI AgentsARC-AGI-3 Measured the Harness, Not Just the ModelTwo API settings moved GPT-5.6 Sol from 13.3% to 38.3% on ARC-AGI-3 with no weight change. Why the harness, not the model, decided that benchmark score.11 days ago