Imported from 54xkeee/mathmod-full-pipeline (
competition/skills/mathmod-reasoning/SKILL.md). Install upstream withnpx skills add 54xkeee/mathmod-full-pipeline --skill mathmod-reasoning. Copyright stays with the author.
MathMod Reasoning
ROLE
Use this Skill for the project's deep reasoning and consequential semantic decisions. ChatGPT is often a strong host for this work, but no particular host is required. This Skill is a set of high-value reminders and context recipes, not a workflow engine or an agent manager.
Start from what the problem means, not from a familiar model name. Clarify the requested action, target, observation, entity/observation unit, time and information availability, constraints, identifiability, assumptions, dependencies, leakage risks, and what evidence could actually validate an answer. Only then compare methods.
You may explore alternatives, challenge your own interpretation, inspect evidence, and load references as needed. Do not save the exploration itself. Publish only durable semantic outcomes that another fresh worker needs to act correctly.
Formal numbers belong to successful formal runs and current evidence. Probe observations may change a decision but are not formal evidence.
READ
Use progressive disclosure rather than loading the whole repository.
- D0: read
AGENTS.mdandCONTEXT.md. - Read the problem statement/rules relevant to the current question.
- Read the smallest relevant D1 semantic homes and any current handoff:
context/model.mdfor interpretation, decisions, assumptions, route, dependencies, validation intent;context/data.mdfor units, entities, observation/time/split semantics, leakage and checked data meaning;context/evidence.mdwhen reviewing existing formal results;context/paper.mdwhen reviewing claims or paper structure;context/visuals.mdonly when visual semantics matter.
- Open D2 artifacts—raw data, source, a formal/probe run, paper text, render-ready data, or authoritative sources—only when the current reasoning needs them.
- Use D3 Git/old artifacts only for conflict, provenance, or understanding what changed.
Reading depth is guidance, not a permission boundary. If correctness requires deeper inspection, inspect it. If a context file is missing, work from the strongest available evidence and state the gap instead of inventing prior decisions.
THINK
The following are reminders, not mandatory checklists. Apply only what is consequential.
Understand the problem before choosing a model
Pay special attention to:
- what each question is actually asking to output;
- requirement vs observation vs decision vs target vs derived quantity;
- entity and observation unit, repeated measures/panels/trajectories/groups;
- time origin, prediction/decision time, horizon, and information available at that time;
- hard/soft constraints and units;
- whether the requested quantity is identifiable or directly observable;
- ambiguity and assumptions that materially change the solution;
- target leakage, future leakage, repeated-entity leakage, split mismatch;
- dependencies between questions and what downstream work becomes stale if a decision changes;
- what validation would support the intended claim and what it would not prove.
A model name is a candidate implementation, not an interpretation of the question. Prefer the simplest route that answers the actual task and can be defended with the data. Do not force novelty, multiple models, sensitivity analysis, or probes when they add no real information.
Before committing a route, connect the problem verb/target, decision variable or unknown, model objective or estimand, validation, and final answer. When a consequential bridge is uncertain, test a boundary or counterexample where validation could pass but the requested guarantee or decision would fail; preserve conditioning events, populations, units/denominators, and output scope. Define decision criteria before model selection; allow constraints, Pareto answers, or parameter sensitivity rather than inventing weights. Separate existence, detectability, and impact, and predictive from decision value when relevant. Disclosing a gap improves honesty but does not complete the missing task. Keep only consequential distinctions in the existing model context; no compulsory extra table, gate, or experiment for a direct, adequate answer.
Use probes only for high-value uncertainty
If an unresolved ambiguity or route choice matters, ask whether a cheap observation can change the decision. A useful probe brief states:
Unknown / competing interpretations
Why the distinction matters
Decision that could change
Minimal experiment or inspection
Discriminating observation
Expected output
What the probe cannot establish
Escalate/stop condition
Reasoning work designs the probe; compute-capable work executes it. The same capable agent may do both. Do not turn probe numbers directly into paper/formal evidence.
Reconsider explicitly when evidence conflicts
When new evidence disagrees with current semantics, compare the existing interpretation and plausible alternatives against the original wording, checked data meaning, constraints, identifiability, and downstream consequences. Prefer the smallest semantic change that explains the evidence. If ambiguity remains, preserve it and state what would resolve it.
Do not silently rewrite target, unit, split, time semantics, constraints, assumptions, route, or paper claims. Identify affected formal runs/figures/text so they can be rerun or revised after the decision changes.
Review independently enough to catch inherited mistakes
For model/result/paper review, first reconstruct the relevant problem meaning from the original statement and verified raw facts before reading current semantic homes or the proposed solution. Then compare the implementation/evidence/claim against that reconstruction.
Check only consequential issues: semantic fit, data support, identifiability, units, constraints, time/split/leakage, implementation invariants, validation, uncertainty when material, extrapolation, feasibility vs optimality, association vs causation, and whether the evidence really supports the claim strength.
Do not invent replacement numbers. Cheap independent scratch checks are allowed under the review reference; substantial recomputation goes to a compute task. Keep formal project content read-only and bind findings to the inspected run/file versions.
TASK LENSES
These names remain for compatibility with the existing fixtures. They are task lenses, not stages, states, owners, or a required sequence:
UNDERSTAND— reconstruct problem/data semantics and validation implications.RECONSIDER— reopen a consequential interpretation after new evidence or conflict.PROBE_DESIGN— design the cheapest experiment that can change a decision.MODEL_REVIEW— independently inspect semantic/model fit and validation logic.RESULT_REVIEW— test whether current formal evidence supports the conclusions.PAPER_SEMANTIC_REVIEW— review technical truth, symbols, claim-to-evidence links, and conclusion strength; leave grammar/layout/visual polish to writing tasks.
Choose the smallest useful lens or combination. The user/task may require work that does not fit a named lens; use judgment rather than inventing another state machine.
KNOWLEDGE: LOAD ON DEMAND
Use these direct routes instead of rediscovering the framework layout:
- problem/semantic understanding → problem understanding and, when data meaning matters, data semantics;
- method selection → method selection → method routing → method library → recipe index → the relevant file under recipes;
- probe design → probe design → the selected recipe only;
- model/result review → review and validation → the relevant model audit → failure patterns;
- source verification → source policy → authority policy, rule policy, and the source manifest.
Open only the smallest route that helps the current question. These references are procedural help, not proof or mandatory model menus. Follow a deeper linked source when correctness requires it. External advice never becomes an official MUST without an authoritative source.
WRITE BACK
The repository is durable cross-session/cross-host memory. Write only information that must survive a fresh session and can change downstream work.
Durable semantic decision
For a consequential choice, use a compact block when it improves recoverability:
Current decision
Scope
Basis
Rejected alternative
Why rejected
Downstream effects
Reopen if
Evidence
Write it to the correct semantic home—normally context/model.md, or context/data.md
for checked data meaning. Do not create a host-specific memory file, a second context
registry, or a transcript archive.
Handoff
Use a handoff only for task-specific information not already stored durably:
Objective
Why this matters
Read first
Current decisions to preserve
Open questions
Expected outputs
Write back
Escalate if
Reference existing sections instead of copying them. For example, preserve
context/model.md#q2-observation rather than restating the entire decision.
Before formal transfer between people/agents/tools, make sure the durable semantic outcome is written to the repository and use a meaningful Git synchronization point when the environment permits.
Before stopping, reconcile CONTEXT.md with the semantic home just updated: remove work
that is now complete from the current focus and confirm its evidence/paper/visual pointers
still describe the current state.
Review return
A review should be compact and actionable:
Verdict: sound / conditionally sound / not sound, with one short basis;Major: correctness-threatening findings only;Minor: limited-scope semantic issues;Evidence: paths/sections/artifacts inspected or missing;Next action: smallest repair and the work type best suited to do it.
Every finding should name where the problem is, why it matters, and the minimal repair. Do not turn review into a release gate.
DO NOT PERSIST
Do not write shared context merely to record activity. Keep these in the live session or their native artifact locations:
- chain-of-thought or full transcripts;
- ordinary brainstorming and abandoned scratch;
- command-by-command history;
- copied datasets/logs/run metadata;
- full paper text duplicated into context;
- conclusions that are already represented by a referenced formal artifact.
BOUNDARIES
- Reasoning decides semantics; it does not create formal numbers.
- Compute-capable work executes data/code/probes/formal experiments and publishes reproducible evidence; it must surface semantic conflict rather than silently reinterpret the task.
- Writing/visual work may improve communication and presentation but must preserve numbers, formulas, units, grouping, variable meaning, model relationships, and claim strength.
- The same person or capable agent may perform multiple or all work types; these are scientific responsibility boundaries, not host permissions.
- Do not add
.mathmod, MCP, scheduler, coordinator, task lifecycle, claim/lease, transaction, release authority, per-host memory, or a context compiler to solve an ordinary Competition reasoning task.