PROACTIVE AGENT EVOLUTION

The Agent discovers it first,
and evolves in my way.

Agentic Shapingis a way of working in which an AI Agent discovers tacit knowledge, preferences, failures, and data in conversations and tasks without waiting for them to be pointed out, shapes them into rules, memories, schemas, templates, validators, and tools, and applies them first on its own in the next execution.

Apply it in 5 minutes

It is not just for coding. The same approach works for documents, analysis, images, video, and deployment.

SIGNALSTone · preferences · corrections · failuresActively detected during work without being explicitly stated

Discover and structure it yourself ↓

AGENT + SYSTEMMemory → Rules → Tools → VerificationRecall and apply it before the next execution
First executionFast, accurate evolution
A scene where an Agent actively organizes human corrections and traces of work into reusable cards and rules
Agentic Shaping does not confine human thought within mechanical frameworks. It means the Agent notices small signals alongside people and refines them into assets that can be retrieved and used in future work. AI-generated visual · 2026

01 · WHY

You got a good result, but
why do you have to explain everything from scratch next time?

When working with AI, you can get fairly good results. But when the conversation ended, I often lost not only why I liked certain choices and what failures I disliked, but also what I considered complete.

So I changed the way I thought about it. Instead of having the Agent only produce results, I wanted it to discover me through the work and change the way it approached future work on its own. This is where Agentic Shaping begins.

PASSIVE AGENTRequest → Result → Forgetting
→
SHAPING AGENTDetect → Structure → Apply → Evolve

02 · START IN 5 MINUTES

That is enough explanation.
Paste this prompt first.

It works best when added to your project instructions. You can also try it in a single conversation first.

1Paste it

Write what you usually need to do today after the prompt.

2Make one correction

Correct it precisely by saying, “No, what I want is…”

3Check the remaining assets

See what remains among the rules, templates, tests, and memories.

4Verify it in the next task

Check whether you need to explain less while maintaining quality.

03 · HOW IT WORKS

Not only the result,
but also the way future results are produced.

The key is a cycle in which feedback becomes an asset for the next task.

A cycle in which human signals pass through structured assets and real verification before returning to the next task
Once a signal is noticed, it becomes the foundation for the next task. The flow is laid out clearly in the steps below.
01DetectRepeated explanations · corrections · failures
→
02CaptureCause · desired direction · scope
→
03Structure when usefulMemory · rules · schemas
→
04ApplyIntegrate into the authoritative path
→
05Verify · simplifyActual evidence · remove duplication

New corrections and results return to the first step ↺

What the AGENT continues to handleMeaning · context · ambiguity · creativity

Judgments that require interpretation, such as “Why is this good?”

What to move into the systemRepetition · format · invariants · execution order

Judgments that can be determined, such as “This condition means failure.”

04 · CONTEXT-SCALABLE ANALYSIS

As things grow, do not make it read more,
make it query more precisely.

As documents and vibe-coding source code accumulate through Agentic Shaping, repeatedly rereading every file in an unstructured way becomes the next bottleneck. Slower exploration, repeated full scans, truncated output, and missed dependencies are signals that the analysis method itself should be structured.

Growing sources and analysis burden Documents · source code · configuration · logs Repeated full scans · long tool output Context pollution · slow exploration · omissions
STRUCTURE →
inventory/ Files · documents · ownership index/ Titles · symbols · reference locations graph/ Calls · dependencies · impact relationships queries/ Scope · filters · evidence bundles gates/ Freshness · consistency · regression
What the analysis tools should guarantee Source tracing · version/hash · refresh on change

Do not turn summaries into a new source of truth. Prove which original source each result came from, and make stale indexes fail instead of using them silently.

How the AGENT will work Existing tools first · small evidence bundles · semantic judgment

If search, parsers, compilers, or tests can answer the question, reuse them first. Build an integrated analyzer only when they are insufficient, and read only the evidence needed for the current question before interpreting it.

This principle is grounded in Lost in the Middle research, which recommends token-efficient tools and curated context context engineering guidance, research showing that repeated search improved repository-level code generation RepoCoder, and work on efficiently updating and querying Tree-sittersyntax trees. The key is not a longer prompt, but a smaller, high-signal analysis surface.

05 · WHERE TASTE LIVES

“My style” does not remain only in words;
it lives in assets that can be found and used at any time.

Choose a single authoritative place that the next Agent can consult before working.

A scene where scattered tacit knowledge is organized into a living work system through active pattern detection
Scattered tacit knowledge becomes a continuously refined working system. The Agent does more than produce results: it identifies recurring judgments and connects them so they can be used in future work.
Signals in the work“This tone sounds too much like a report.”“Don’t just check deployment success—look at the actual URL.”“Catch this error earlier next time.”
SHAPE →
memory/ Preferences · judgmentsrules/ Instructions · rubricsschemas/ Types · contractstests/ Failures · fixturetools/ Commands · validators

06 · MEMORY MAKES IT STRONGER

The more memory carries forward,
the better Agentic Shaping works.

You can start without a memory tool. But when the corrections and decision criteria noticed during work can be found again before the next task, the Agent can finally align consistently with the user’s way of working.

Start right away Agentic Shaping

Find and apply signals within the current conversation and project instructions. If the conversation ends, the context built up along the way may disappear with it.

Works better + LLM Wiki

Keep corrections, preferences, and decision criteria over time and retrieve them before the next task. That is why Agentic Shaping works better.

Memory is not complete when it is merely collected. It becomes powerful when it is found before the next task and applied to the actual plan and result.

07 · COPY & RUN

When you are stuck,
pick one and use it right away.

You do not need to build a massive system from the start.

Start a task
Complete this task. First find relevant memories and project rules, then establish success criteria and actual verification. Capture recurring judgments, corrections and failures as improvement candidates, but choose the simplest suitable path among existing-tool reuse, direct model processing and reusable code by total cost and required accuracy. Create new assets only when beneficial; do not create new code or a cost assessment form for every task. Preserve mandatory verification and distinguish verified improvements from effects that have not been measured. When structured checks cannot decide, first read the relevant original material, code and failure output directly to identify the problem and next action.
Problem recurrence
Diagnose this problem through Agentic Shaping. If structured checks cannot identify the cause, first directly read the relevant original material, code and failure output to identify the problem and next action. Compare existing-tool reuse, direct handling and necessary code repair to fix the cause; implement prevention assets only when beneficial. Preserve required verification, then remove duplicate and temporary paths after actual checks.
Retrospect after the task
Review the task just completed from an Agentic Shaping perspective. Distinguish ① my preferences and decision criteria you should have detected on your own, ② repeated manual judgments, ③ failures discovered late, ④ reusable assets already created, and ⑤ the next candidates for structuring. Apply only what has long-term value to the authoritative repository with its scope and evidence, and recall it before the next task.

08 · REAL EXAMPLES

One correction
becomes the default for the next execution.

DOCUMENT

“Give me the copyable prompt and the starting order first, not a conceptual explanation.”
→ The order of 5-minute start → copy prompt → inputs/outputs → examples → concepts remains as a documentation rule.

CODE

“The Agent had to find the same version difference manually again.”
→ A version manifest + compatibility check + diagnostics + regression tests catch it before the next update.

ANALYSIS

“The conclusion is correct, but the reasoning path cannot be reproduced.”
→ An input schema + decision rubric + evidence + calculation fixture preserves the path to the conclusion.

SCALE

“As the repository grows, the Agent reads everything every time and gets slower.”
→ An inventory tracking source locations and versions + a symbol/dependency graph + scoped queries provide only the necessary evidence in small bundles.

MEDIA

“My writing style and screen style disappeared again.”
→ Style memory + good/bad examples + rendering checks are applied before generation.

09 · MEASURE

The words “it is evolving”
are proven in the next execution.

Re-explanation ↓

Manual judgment ↓

Late failures and retries ↓

Time and cost ↓

Full reanalysis and context usage ↓

Intent alignment and reproducibility ↑

10 · BEHAVIORAL VALIDATION

Does it really trigger?
We compared them using the same task.

We did not mix general task capability with the triggering of Agentic Shaping. We gave both groups the same current task and included the exact public prompt only in the treatment group. With current-task completion and safety protected by separate gates, we scored whether signal capture, formalization, and improvement of the next execution were actually added.

AGENTIC SHAPING ACTIVATION · 26 PAIRED SCENARIOS · 52 AGENT RUNS 61.5% → 96.2%

Scenarios that performed every distinctive behavior were 16/26 → 25/26. Detailed behaviors were 82/105 (78.1%) → 104/105 (99.0%), with the treatment group ahead in 10 · tied in 16 · behind in 0. The five completely unseen final holdouts also improved from 3/5 (60%) → 5/5 (100%).

CURRENT TASK & SAFETY GUARD 27/27 ↔ 27/27

Both the baseline and treatment groups selected all current-request completion criteria, and made 0 selections across 52 prohibited behaviors. We did not sacrifice current-task completion or safety to widen the score gap.

REAL FILE EXECUTION · 1 TASK · 6 CHECKS 6/6 ↔ 6/6

This is a regression gate that actually modified a broken version-selection repository. Because both conditions passed this task, we use it only as evidence that there was no actual file-execution regression—not as evidence of prompt superiority.

CONTEXT-SCALABLE ANALYSIS 1/5 → 5/5

From full rereads to measurement, integration with existing tools, incremental graphs, and freshness gates.

TEMPORARY SCRIPT 0/4 → 4/4

From the current conversion to a common contract, canonical commands, regression tests, and deduplication after verification.

FINAL HOLDOUT · CONFIG DRIFT 1/4 → 4/4

From a single configuration change to type contracts, preflight checks at every entry point, and fixtures.

VALIDATION SHAPES TOO

Not just the prompt,
the verification system evolves too.

This verification was not complete from the start. We captured small samples, target behaviors mixed into the baseline group, ambiguous cases, repeated single-case failures, and long-running execution timeouts as shaping signals. We promoted them into 26 discriminating scenarios; separate task and safety gates; development/holdout separation; an early filter for failure cases; preservation of repeated results; and hash-matched checkpoints, resume, and timeout retries. In the end, tests/ we even found that the evaluator could not locate the passing tests below and fixed it with recursive discovery.

  1. DetectDetection of inflated percentages, evaluator contamination, and repeated failures
  2. StructureSeparate discriminating scores from general quality regressions
  3. Verify earlyReplay only failure cases first, then run the full regression
  4. IntegrateIdentical-hash checkpoints and reproducible result assets
CONFIRMED FINAL POLICY SYNC 90% → 100% HOMEPAGE v0.6 ↔ SLOGS 2026.10.05.2

We synchronized the confirmed behavioral contract from the homepage’s final prompt with the Korean and English Slogs LLM Wiki runtime policies. Across five live-policy pair regressions and 10 Luna Max runs, distinctive behaviors improved from 18/20 → 20/20, with 0 prohibited behaviors.

  • Independently execute current-result completion · safety boundaries · durable improvements
  • Do not force memories or global rules onto one-off tasks
  • Complete selected structuring through authoritative assets, early checks and regression
  • Consolidate repeated versions, paths, and configurations into a single authority and replace all related hardcoding
Policy behavior results Confirm the final Korean policy →
REUSABLE SKILL DISCOVERY · AS-SA-001 25/25 BEHAVIOR VERIFIED User-facing changes

Classify signals as local · project · cross-project · general-method, and synthesize only the latter two levels into de-identified general-purpose skills. Block personal information, project confidential information, credentials, secrets, insufficient generalization, and failures in normal, boundary, or negative evaluations before registration.

  • Technical notes
  • Agentic Shaping abstraction and safety checks: 25/25 passed
  • Slogs natural-language discovery/registry 36/36 · PostgreSQL 1/1 · 251 total passed · 0 failed
  • Candidates are excluded from search, selection, and application before review
  • Awaiting initial scope selection · Windows validation · External locator rehashing unsupported
AS-SA-001 contract Deterministic evaluation cases →
Sample and score definitions

26 paired scenarios · 52 Agent runs · 105 distinctive behaviors · percentages are the proportion of scenarios with a complete pass in the fixed set

Execution environment

GPT-5.6 Luna · Max · Codex CLI 0.149.0-alpha.4.3 · isolated execution · 2026-08-25

Reproduce directly

How to run · Discriminating results · File execution results · Evaluation design notes

Verification conclusion: In a single Luna Max run on this fixed set, Agentic Shaping’s distinctive behaviors increased clearly, with no regressions in current-task completion or safety. The treatment group also missed one behavior: “replace all hardcoding” for version drift. These figures are not population estimates or guarantees across all model and tool environments. We therefore continue shaping the prompt and evaluator together while publishing the numerators, denominators, ties, omissions, and failure history.

11 · ORIGIN & DIFFERENCE

From Compound Engineering
This is not just a name change.

Compound Engineering’s execution structure—“one task makes the next task easier”—was a valuable reference. In particular, its approach of exposing planning, work, review, and accumulation through commands and files directly informed the rebuilding of this guide.

Agentic Shaping starts with what emerges between people and Agents: tacit knowledge and corrections. An Agent does not remain a tool that waits for instructions; it first discovers signals in the work, turns them into structures reusable across coding, documentation, analysis, and media, and improves the next execution itself.

Vibe Compiler made formalization look like the goal, while Vibe Tailoring emphasized only customized results and weakened the Agent’s agency. The name Agentic Shaping captures a working system refined together with an Agent that acts on its own.

Try just one thing today.

Take one sentence you keep explaining
and have the Agent notice it first and preserve it as a rule for next time.

Go copy the starter prompt ↑
Prompt copied.