GPT-6 Astra
1,050,000token context window
A larger window doesn't ensure the agent has read the right files. OpenAI's Codex guidance recommends supplying task-relevant files and examples.
Help your coding agent keep track
of a multi-file change.
research.mdplan.mdchanges.mdreview.md1,050,000token context window
A larger window doesn't ensure the agent has read the right files. OpenAI's Codex guidance recommends supplying task-relevant files and examples.
1Mtoken context window
Anthropic describes Opus 5.5 as a model for long coding sessions. Its context guidance still warns that longer inputs can reduce recall.
Provider API limits. Effective limits depend on the client and configuration. GitHub lists both models as generally available in Copilot as of September 23.
The next response uses the context the client sends. That can include earlier messages, file reads and tool results, or summaries of older turns.
A search returns 100 matches. The relevant module is in the results that weren't shown.
RiskThe agent misses an exception
Failed attempts and long logs remain after they've stopped helping with the task.
RiskImportant details are harder to use
A nearby module uses a legacy convention. The agent copies it without checking whether it still applies.
RiskThe change follows the wrong convention
query: "variable" includePattern: "modules/**"Found 312 matches in 14 files for "variable"
(showing 100 matches in 14 files)
modules/blob/variables.tf
3:variable "resource_prefix" {
... (19 more matches in this file)
modules/legacy-table/main.tf[File content truncated at line 2000. Use read_file
with offset/limit parameters to view more.]
Search each module or read the next range before treating this as complete coverage.
Fictional repository; VS Code 1.139.0 behavior. Search defaults to 100 results with a 200-result cap, both experiment-based settings. Other hosts use different limits.
Old reads and logs can compete with the details needed for the next step.
Across 18 models, longer inputs often reduced accuracy on controlled tasks.
Irrelevant sentences reduced accuracy on the math problems tested.
Models used facts in the middle of long inputs less reliably.
Bars are conceptual, not measurements. The studies tested earlier models, not the two models shown earlier in this talk.
If you only need one phase, run its skill directly:
/rpi-research/rpi-plan/rpi-implement/rpi-reviewRPI Agent is optional and uses the same phase skills. This Chat example doesn't send a request.
research.mdInvestigates without changing source code.
plan.md + critiqueDefines tasks and checks; gets one critique.
changes.mdChanges the code and records check results.
review.mdCompares the result with the requirements.
### Variable names differ across modules
12 of 14 modules use resource_prefix.
2 legacy modules use prefix.
Evidence: C1, C2
Coverage: all 14 variables.tf files read.
First search: 100 of 312 matches shown.
Follow-up: searched each module.
Limits: only variable declarations checked.
## Planning Readiness and Next Step
Not ready: decide whether to keep legacy names.
Illustrative findings and evidence IDs. Subagents are optional in every RPI phase as of September 18 (#2902).
#### [ ] P01-T01: Add log retention
Goals:
* All 14 modules accept log_retention_days.
Requirements:
* FR-001: Default to 30 days.
* FR-002: Accept only 1 to 365 days.
* FR-003: Keep prefix in legacy modules (D1).
Details:
* Follow the pattern in modules/blob.
References:
* Research C1, C2; decision D1
Dependencies:
* None
Put the goal, requirements and source references together, so the implementer can find them without retracing the investigation.
Here, a value of 0 must fail in every module, including the two legacy modules.
rpi-plan-critique assesses the final plan once: Pass, Revise or Blocked.
research.mdplan.md + critiquechanges.mdreview.mdUse /clear or a new chat when old work or repeated corrections get in the way.
Reference the current plan and supporting files. Name the task, such as P01-T01.
/compact summarizes the chat and can omit details. Save important decisions in the files first.
An optional way to work. RPI records live in dated folders under .copilot-tracking/, which is gitignored in HVE Core.
changes.md.rpi-implementrpi-planrpi-researchfollow-up itemExample statuses. Complete means the review finished. The outcome tells you whether the result meets the plan.
Example selection. Safety confirmations, required gates and human review apply in every mode.
You can also use one phase skill, such as /rpi-research, without the full workflow.
A written requirement can still be wrong.
Research and review take time. Use them where they help.
Read the findings and plan before approving the change.
Copyright (C) 2011-2026 Hakim El Hattab, http://hakim.se, and reveal.js contributors Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software. THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.