fix(harness): resolve layer contradictions + context-budget caps for airouter TDD
- tdd-gate: pass criterion is the task's TEST_FILES scope (matches run-ci-tests.sh M19 P-14), not the full suite — contradicted qwen-code-methodology rule 5 in the same system prompt - tdd-protect: tests/ path described relative to project working dir (AGENT_WORKING_DIR is /workspace/project in repo composites) - qwen-code-methodology: task_complete tool DOES exist — rule 8 said the opposite; pytest examples use -q - airouter-code-tdd template: TT-39 runtime_overrides (max_result_bytes 24KB, max_history_turns 16) Claude-Session: https://claude.ai/code/session_012arAwUuBQrpnMzJ36mEVBk
This commit is contained in:
@@ -15,11 +15,11 @@ You are running on **Qwen3.6-27B** via the agentic tool-calling runner. Thinking
|
||||
If this shows deleted non-blank lines from existing code, you have broken something — revert and try again with a targeted `Edit`.
|
||||
5. **Only run the specific test file for your change.** Never run the full test suite.
|
||||
```
|
||||
python -m pytest tests/test_<module>.py -v --tb=short -x
|
||||
python -m pytest tests/test_<module>.py -q --tb=short -x
|
||||
```
|
||||
6. **Do not create backup copies** (`*_orig`, `*_old`, `*_bak`, `*_backup`).
|
||||
7. **Do not rename existing files before modifying them.**
|
||||
8. **When the task is done, respond with plain text and stop.** Do not call any tool to signal completion. There is no "finish", "done", or "report" tool — emitting one wastes a turn and the runner will treat it as more work.
|
||||
8. **When the task is done, call the `task_complete` tool once** (with a one-paragraph summary), or respond with plain text and stop. Do not invent other completion tools.
|
||||
|
||||
## How to work
|
||||
|
||||
|
||||
@@ -4,23 +4,23 @@ You are operating in TDD mode. The test suite defines correctness. **You must no
|
||||
|
||||
## Rules
|
||||
|
||||
1. **Run the full test suite before committing.** Every change must be followed by a test run.
|
||||
1. **Run the task's test scope after every change.** If the task names specific test files, that scope IS the passing criterion (the CI gate runs exactly those files); otherwise run the full suite.
|
||||
2. **A red test suite is not acceptable output.** If tests fail, keep working.
|
||||
3. **Never skip, comment out, or delete tests.** Tests are read-only (OS-enforced). Any attempt to modify them will fail.
|
||||
4. **Never use `pytest --ignore`, `-k "not ..."`, or `--deselect` to hide failures.** The passing criterion is `pytest` with no filters.
|
||||
4. **Never use `pytest --ignore`, `-k "not ..."`, or `--deselect` to hide failures.** The passing criterion is `pytest` on the task's test scope with no filters.
|
||||
5. **Fix the implementation, not the tests.** If a test looks wrong, you are wrong.
|
||||
6. **Commit only when green.** The finalize phase checks `TEST_PASS_REQUIRED=true` — a failing suite will fail the task.
|
||||
|
||||
## Workflow
|
||||
|
||||
```bash
|
||||
# Run tests to see current state
|
||||
python -m pytest tests/ -v --tb=short
|
||||
# Run the task's test scope to see current state (use the file(s) named in the task)
|
||||
python -m pytest <task test file(s)> -q --tb=short
|
||||
|
||||
# Implement until green
|
||||
# ... edit source files ...
|
||||
|
||||
python -m pytest tests/ -v --tb=short
|
||||
python -m pytest <task test file(s)> -q --tb=short
|
||||
|
||||
# Only when all green:
|
||||
# finalize will commit and push via agent-repo
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
## The tests directory is READ-ONLY
|
||||
|
||||
`/workspace/tests/` has been locked at the OS level. You **cannot** write to it. Any attempt will fail with a permission error.
|
||||
The `tests/` directory of the project working directory has been locked at the OS level. You **cannot** write to it. Any attempt will fail with a permission error.
|
||||
|
||||
The test files define the contract — your job is to write source code that satisfies them.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user