Loading the presentation. If it does not open, download the complete HTML file and open it in a browser with JavaScript enabled.
MICROSOFT / HVE CORE / SEPTEMBER 2026

RPI with
HVE Core

Help your coding agent keep track
of a multi-file change.

SOURCE SNAPSHOT / SEPTEMBER 23, 2026
  1. 01
    ResearchFindings and open questions
    research.md
  2. 02
    PlanTasks and required checks
    plan.md
  3. 03
    ImplementCode changes and results
    changes.md
  4. 04
    ReviewFindings and follow-up
    review.md
microsoft/hve-coreIncludes scripted examples
01 / CONTEXT

Larger windows still need good context

OPENAI / SEPTEMBER 3, 2026

GPT-6 Astra

1,050,000token context window

A larger window doesn't ensure the agent has read the right files. OpenAI's Codex guidance recommends supplying task-relevant files and examples.

ANTHROPIC / SEPTEMBER 22, 2026

Claude Opus 5.5

1Mtoken context window

Anthropic describes Opus 5.5 as a model for long coding sessions. Its context guidance still warns that longer inputs can reduce recall.

A model can hold more information and still miss the detail your task depends on.

Provider API limits. Effective limits depend on the client and configuration. GitHub lists both models as generally available in Copilot as of September 23.

CONTEXT / WHAT THE MODEL RECEIVES

Your prompt shares space with the session

The prompt is only part of the input

The next response uses the context the client sends. That can include earlier messages, file reads and tool results, or summaries of older turns.

  • Old instructions can remain after the task changes.
  • Large tool outputs can take up much more space than your request.
  • When the window fills, VS Code summarizes older conversation history.
What can be in the context
  • System prompt and tool definitions
  • Your request and attached files
  • Instructions and skills
  • Files the agent read
  • Search and tool results
  • Terminal output
  • Results from connected tools and subagents
  • Summaries of older turns
Input for the next action
CONTEXT / FAILURE MODES

How useful context gets lost

01 / MISSING

Some evidence is missing

A search returns 100 matches. The relevant module is in the results that weren't shown.

RiskThe agent misses an exception

02 / NOISE

The chat fills with old work

Failed attempts and long logs remain after they've stopped helping with the task.

RiskImportant details are harder to use

03 / DISTRACTION

An old pattern looks reusable

A nearby module uses a legacy convention. The agent copies it without checking whether it still applies.

RiskThe change follows the wrong convention

When an answer is wrong, check what the agent saw as well as what it did.

MISSING EVIDENCE / VS CODE 1.139.0

A search result may be incomplete

Agent tool calls Reconstructed from source
grep_searchquery: "variable" includePattern: "modules/**"
Found 312 matches in 14 files for "variable"
(showing 100 matches in 14 files)

modules/blob/variables.tf
3:variable "resource_prefix" {
... (19 more matches in this file)
read_filemodules/legacy-table/main.tf
[File content truncated at line 2000. Use read_file
with offset/limit parameters to view more.]

Read the limits with the result

  1. Only 100 of the 312 matches are shown.
  2. The file read ends at line 2,000.
  3. The agent needs another call to read the rest.

Search each module or read the next range before treating this as complete coverage.

Fictional repository; VS Code 1.139.0 behavior. Search defaults to 100 results with a 200-result cap, both experiment-based settings. Other hosts use different limits.

NOISE AND DISTRACTION

A longer chat can make details harder to find

Illustration of context use, not measured proportions. The first bar has room for more evidence. In the second, old reads, failed attempts and logs use much of that space, leaving less room for new information.
Start of a task
  • Instructions and tools
  • Relevant evidence
  • Free space
Later in the session
  • Instructions and tools
  • Relevant evidence
  • Stale reads, failed attempts, logs
  • Free

Old reads and logs can compete with the details needed for the next step.

CHROMA / 2025

Across 18 models, longer inputs often reduced accuracy on controlled tasks.

SHI ET AL. / 2023

Irrelevant sentences reduced accuracy on the math problems tested.

LIU ET AL. / 2023

Models used facts in the middle of long inputs less reliably.

Bars are conceptual, not measurements. The studies tested earlier models, not the two models shown earlier in this talk.

CONTEXT / SCRIPTED VS CODE WALKTHROUGH

How a simple repo question goes wrong

02 / RPI

Carry the findings
into the next phase

RPI records what you've learned and decided.
A new chat can pick up from those files.

Context engineering means choosing the information an agent needs for its next step.

Summary of the idea discussed by Andrej Karpathy and Tobi Lutke, June 2025
RPI / CHAT SETUP

Start with RPI Agent or a phase skill

Let RPI Agent coordinate the work

  1. Open the agent dropdown below the Chat input.
  2. Select RPI Agent.
  3. Describe the task, its constraints and how you'll check the result.

If you only need one phase, run its skill directly:

/rpi-research/rpi-plan/rpi-implement/rpi-review

RPI Agent is optional and uses the same phase skills. This Chat example doesn't send a request.

RPI / WHAT EACH PHASE READS AND WRITES

Each phase has a job and a written result

01

Research

Reads
Task, code and sources
Writes
research.md

Investigates without changing source code.

02

Plan

Reads
Research findings and decisions
Writes
plan.md + critique

Defines tasks and checks; gets one critique.

03

Implement

Reads
Plan tasks and their references
Writes
Source + changes.md

Changes the code and records check results.

04

Review

Reads
Plan, critique, changes and checks
Writes
review.md

Compares the result with the requirements.

Use the files to carry decisions forward. Research runs when evidence is missing; adequate existing evidence can be reused.

RESEARCH

Record what you checked and what is missing

log-retention-research.md Illustrative excerpt
### Variable names differ across modules
12 of 14 modules use resource_prefix.
2 legacy modules use prefix.
Evidence: C1, C2
Coverage: all 14 variables.tf files read.
  First search: 100 of 312 matches shown.
  Follow-up: searched each module.
Limits: only variable declarations checked.

## Planning Readiness and Next Step
Not ready: decide whether to keep legacy names.

Make the evidence easy to check

  • Record which files and sources you examined.
  • Link findings to C# codebase or W# external evidence. State what remains unknown.
  • Check a helper's suggestions against the original source before using them.
WiderDeeperContrarian

Illustrative findings and evidence IDs. Subagents are optional in every RPI phase as of September 18 (#2902).

PLAN / CRITIQUE

Write tasks you can verify

log-retention-plan.md Illustrative excerpt
#### [ ] P01-T01: Add log retention
Goals:
* All 14 modules accept log_retention_days.
Requirements:
* FR-001: Default to 30 days.
* FR-002: Accept only 1 to 365 days.
* FR-003: Keep prefix in legacy modules (D1).
Details:
* Follow the pattern in modules/blob.
References:
* Research C1, C2; decision D1
Dependencies:
* None

Keep the details with the task

Put the goal, requirements and source references together, so the implementer can find them without retracing the investigation.

Here, a value of 0 must fail in every module, including the two legacy modules.

rpi-plan-critique assesses the final plan once: Pass, Revise or Blocked.

CONTEXT BETWEEN PHASES

Start a fresh chat without starting over

  1. Research chatStarts from the taskresearch.md
  2. Plan chatStarts from research.mdplan.md + critique
  3. Implement chatStarts from the planchanges.md
  4. Review chatReads plan, critique and resultsreview.md

Leave unhelpful history behind

Use /clear or a new chat when old work or repeated corrections get in the way.

Bring the relevant files

Reference the current plan and supporting files. Name the task, such as P01-T01.

Check the summary

/compact summarizes the chat and can omit details. Save important decisions in the files first.

An optional way to work. RPI records live in dated folders under .copilot-tracking/, which is gitignored in HVE Core.

IMPLEMENT / REVIEW

Check the change against the plan

IMPLEMENT
  • Follow approved tasks in dependency order.
  • Mark a task done only when its requirements hold.
  • Record the changes and check results in changes.md.
  • Update the plan when a discovery changes the work. Pause tasks that depend on it.
REVIEW / ONE RECORD PER RPI TASK
  • Code misses a requirementrpi-implement
  • A decision is missingrpi-plan
  • Evidence is missingrpi-research
  • Work remains outside scopefollow-up item
Execution: CompleteOutcome: Defects found

Example statuses. Complete means the review finished. The outcome tells you whether the result meets the plan.

OPTIONAL RPI AGENT / PARTICIPATION

Decide how involved you want to be

Example selection. Safety confirmations, required gates and human review apply in every mode.

RPI / SCRIPTED WALKTHROUGH

Walkthrough: add a retention setting

TRADE-OFFS

Use RPI when the work needs investigation

USE RPI
  • The change spans files or modules you don't know well.
  • Different modules follow different conventions.
  • You need a record of decisions or dependencies.
  • Someone else will continue or review the work.
USE A DIRECT EDIT
  • The change is clear and isolated.
  • You could describe the diff in one sentence.
  • Existing evidence and checks are enough.

You can also use one phase skill, such as /rpi-research, without the full workflow.

Check the assumptions

A written requirement can still be wrong.

Allow for the extra work

Research and review take time. Use them where they help.

Human review still matters

Read the findings and plan before approving the change.

GETTING STARTED

Try RPI on
your project

RPI Agent/rpi-research

Give it the task and the constraints.
Read the findings before you approve the plan.