Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
43 changes: 43 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -367,6 +367,49 @@ The three Phase 7 blockers are unchanged; the newest name is `asset_configure_im

### Fixed

- **The live harness copies only the fixture it tracks (#1036).** It copied
`tests/godot_smoke` whole, including the `.didi/` runtime state the Python
suites leave there, so a local run started from that state, and a suite
running at the same time held a lock that failed the copy before the first
assertion. The copy now leaves out `.didi/` and `.godot/`, which a CI
checkout never has.
- **The seven blackboard writers answer with what they saved (#1019, in
part).** `blackboard_write`, `blackboard_patch`, `blackboard_clear` and the
four task calls built their answers from the board in memory before it was
saved, and never read the saved board back. Each now reads the file back
after the save and answers from it; `blackboard_write` also returns the
`value` now stored. They move from exempt to observed in
`tests/observed_post_state.json`, and the live harness compares each answer
with the board file as Godot parses it. Fifteen tools remain on #1019.
- **The local asan lane is no longer killed for memory (#1050).**
`tools/localci` built every lane with Ninja's default of CPUs + 2 jobs, and
ASan builds of the largest files need several GB each, so on a 16 GB Docker
VM the compiler was killed with a message that reads like a compile error.
The asan lane now runs one job per 2 GB of the container's memory, the lane
prints the count it chose, and `DIDI_LOCALCI_JOBS` overrides it.
- **A field trial keeps its evidence and pins its tester (#1005, #1007).**
`trial.py` scored a Claude tester's transcript where the client filed it and
never copied it, so trial 05 can no longer be rescored. It now copies it into
the trial directory as `transcript.jsonl` and scores the copy. And its
`--allowed-tools` list restricted nothing under `bypassPermissions`: trial
06's tester also used `PowerShell`, and earlier ones met the host's skills.
The tester now runs with `--tools` naming seven built-in tools (including
`ToolSearch`, which Didi's deferred tools need), skills off, and only project
settings, and `trial.json` records that as `tester_environment`.
- **Engine output and helper answers carry text, not terminal noise (#1028).**
`runtime_read_output` relayed Godot's colour escape codes in its records, 40
of a fresh headless editor's first 41, so a match against the visible text
failed. A record's message now has them taken out when it is captured. And
the offline helper Godot logged Didi's own INFO line before anything else,
so `shader_check_compile`'s `raw_output` began with it; that line is DEBUG
now, and the answer starts with the engine's own output on 4.5, 4.6 and 4.7.
- **The advice for a new autoload no longer offers a rescan that does not work
(#1002).** `signal_connect`'s `target_script_not_compiled` note and
`LLM_INSTRUCTIONS` said to restart the editor or call
`editor_reload_project`. `project_set_autoload` says the rescan does not
register the singleton, and that is what the engine does: called after
`project_set_autoload` on 4.5.1, 4.6.2 and 4.7.2, the script still does not
compile. Both now say to restart.
- **Two failures that skipped the error floor now name their fix (#1043).** A
confirmation the person declined or dismissed answered `403` with no code
and no remedy, so it read as any other 403. It now carries `data.code:
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@
[![CI](https://github.com/saworbit/didi/actions/workflows/ci.yml/badge.svg)](https://github.com/saworbit/didi/actions/workflows/ci.yml)
[![CodeQL](https://github.com/saworbit/didi/actions/workflows/codeql.yml/badge.svg)](https://github.com/saworbit/didi/actions/workflows/codeql.yml)
[![OpenSSF Scorecard](https://api.scorecard.dev/projects/github.com/saworbit/didi/badge)](https://scorecard.dev/viewer/?uri=github.com/saworbit/didi)
[![Tests](https://img.shields.io/badge/tests-1478-2ea043?logo=pytest&logoColor=white)](docs/TEST_INVENTORY.md)
[![Tests](https://img.shields.io/badge/tests-1486-2ea043?logo=pytest&logoColor=white)](docs/TEST_INVENTORY.md)
[![Release](https://img.shields.io/github/v/release/saworbit/didi?logo=github&color=blue)](https://github.com/saworbit/didi/releases/latest)
[![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](https://opensource.org/licenses/MIT)
[![Godot Engine](https://img.shields.io/badge/Godot-4.5%2B-478cbf?logo=godotengine&logoColor=white)](https://godotengine.org/)
Expand Down
4 changes: 2 additions & 2 deletions docs/FIELD_TRIAL_DESIGN.md
Original file line number Diff line number Diff line change
Expand Up @@ -140,9 +140,9 @@ Three things it does deliberately.

**It scores the bridge, not just the server.** This is the finding trial 03 paid for. A trial is seeded against a server binary and scored against that binary's manifest, while the live half of every call is served by a different file that the tester installs by hand and no artifact recorded. That run spent about an hour concluding a shipped capability did not exist, because the GDExtension answering it was six days older than the server. The seed now records the server's build id, read from `didi --version`, and the SHA-256 of the addon the tester is meant to install alongside the one lying in the repository's own `addons/didi`, which is gitignored, written by no build step, and indistinguishable by eye. After the run, `bridge.py` reads the pairing the server reported on every live session call and returns one of three verdicts. `matched` and `mismatched` are the obvious two. The third is `not_observed`, and it is not a pass: a run in which no call ever reported a pairing does not say which build served it, and its live findings are unaccounted for rather than clean.

**It never trusts the tester's account of itself.** Coverage comes from the transcript, as in section 8, and so does the bridge verdict. The ledger is read by a human afterwards and is not an input to any number here.
**It never trusts the tester's account of itself.** Coverage comes from the transcript, as in section 8, and so does the bridge verdict. The ledger is read by a human afterwards and is not an input to any number here. A Claude tester's transcript is copied into the trial directory as `transcript.jsonl` and scored from there, so the evidence stays with the trial after the client's own store loses it (#1005).

**It can hand the seed to a different tester.** `--engine claude` and `--engine codex` run the same seed and the same briefing under different clients, and the summary names which one so nobody has to infer it later. Trials 01 through 03 were all Claude, and across three runs they agreed that an agent reaches for about 40% of the implemented surface. That figure is either a fact about agents or a habit of one client, and there is no way to tell those apart from three runs of the same client. Trial 04 is the first run by another engine and exists to separate them.
**It can hand the seed to a different tester.** `--engine claude` and `--engine codex` run the same seed and the same briefing under different clients, and the summary names which one so nobody has to infer it later. Trials 01 through 03 were all Claude, and across three runs they agreed that an agent reaches for about 40% of the implemented surface. That figure is either a fact about agents or a habit of one client, and there is no way to tell those apart from three runs of the same client. Trial 04 is the first run by another engine and exists to separate them. A Claude tester's environment is pinned rather than inherited from the machine: seven built-in tools, `ToolSearch` among them because Didi's tools arrive deferred, skills off, and only the project's own settings loaded. `trial.json` records it as `tester_environment` (#1007).

Two differences to carry into any comparison. Codex configures MCP servers in `config.toml` rather than from a file, so the seed's `.mcp.json` is translated into `-c` overrides and one seed artifact still describes the server under test for both engines; `--ignore-user-config` is this engine's `--strict-mcp-config`, and it discards the model along with everything else, which is why the model is named on the command line. And Codex has no equivalent of `--max-budget-usd`, so on that engine the timeout is the only ceiling on a run that has stopped finishing. That is worth knowing before walking away from one rather than after.

Expand Down
2 changes: 1 addition & 1 deletion docs/LLM_INSTRUCTIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -103,7 +103,7 @@ Nodes inside an instanced sub-scene belong to that scene's file. `scene_get_hier
- Use typed autoload and InputMap tools for `autoload/*` and `input/*`; never route those namespaces through `project_set_setting`.
- To localise a game, register the `.translation` files in `internationalization/locale/translations`, never the `.csv` they came from. Godot's importer writes them beside the CSV as `<name>.<locale>.translation` and lists them under `dest_files` in the CSV's `.import` sidecar. `project_set_setting` refuses the `.csv` with `retry_with` holding the list with the CSV replaced by those files: resend with it. It also refuses a missing path, a `.tres` and any other file Godot does not register a translation from (#989). A `.po`, `.mo` or `.res` Translation registers too. Set `internationalization/locale/test` to a locale such as `fr` to make a launched game use it.
- `project_set_setting` is the only project writer that works with no session attached. Without one it edits `project.godot` and reports `execution_mode: "offline_fallback"`. Use it to enable the addon in a project that does not have it, by setting `editor_plugins/enabled` to an array containing `res://addons/didi/plugin.cfg`, after the addon folder is in place. Nothing has loaded the value when the call returns; start Godot and attach.
- `project_set_autoload` persists the setting but cannot make the attached editor register the singleton; it returns `requires_editor_restart: true`. Until that editor restarts, every script referencing the new singleton fails to compile with `Identifier not found`. Those errors are the tool's doing, not your script's, so do not rewrite working code to chase them. `signal_connect` refuses a handler in such a script with `409 target_script_not_compiled`, naming the script and the singleton it cannot resolve, rather than reporting the method absent. Restart the editor or call `editor_reload_project`; do not rename the handler.
- `project_set_autoload` persists the setting but cannot make the attached editor register the singleton; it returns `requires_editor_restart: true`. Until that editor restarts, every script referencing the new singleton fails to compile with `Identifier not found`. Those errors are the tool's doing, not your script's, so do not rewrite working code to chase them. `signal_connect` refuses a handler in such a script with `409 target_script_not_compiled`, naming the script and the singleton it cannot resolve, rather than reporting the method absent. Restart the editor; `editor_reload_project` is a file rescan and does not register the singleton. Do not rename the handler.
- `script_check_syntax` reports autoload identifiers at `severity: "warning"` with a `note`, not as errors. The Godot compiler check runs in a process with no SceneTree, so it can never see an autoload; that is permanent and separate from the restart limitation above. `has_errors` is therefore a verdict about your script. Do not rewrite working code because a warning names an autoload.
- `script_check_syntax` only asks the Godot compiler when you give it a `file_path`. Given `source_text` there is no file to compile, so it runs Didi's own lexical rules alone and `has_errors: false` means "no bracket, indentation or tab-and-space problem", not "this compiles". Read `engine_checked` on every answer: `false` means no compiler was asked, and the answer carries a `limitation` saying what it covers. To check a draft properly, send it to `project_verify_changes`, which compiles it in an isolated copy of the project, or write it with `script_create` and check the file.
- Treat `replace: true`, `overwrite: true`, and `discard_unsaved: true` as explicit destructive intent. Do not add them speculatively.
Expand Down
18 changes: 9 additions & 9 deletions docs/TEST_INVENTORY.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,22 +10,22 @@ Every number here is derived from the suites themselves rather than written down

| Measure | Count |
| --- | ---: |
| Automated tests | **1478** |
| Live-harness assertions | 1386 |
| Automated tests | **1486** |
| Live-harness assertions | 1388 |

The badge in the [README](../README.md) shows the automated test total. Harness assertions are counted separately because they are assertions inside one live scenario, not independently runnable cases; adding them together would flatter the number.

## Native C++ suite (`didi_tests`)

**944 tests.** Derived from `didi_tests --list`, which prints the registry the runner iterates.
**946 tests.** Derived from `didi_tests --list`, which prints the registry the runner iterates.

| Suite | Tests |
| --- | ---: |
| `AnimAddLibrary` | 3 |
| `AssetConfigureImport` | 6 |
| `AudioAddBus` | 10 |
| `Base64` | 1 |
| `Blackboard` | 13 |
| `Blackboard` | 14 |
| `BlackboardResources` | 11 |
| `BlackboardTasks` | 8 |
| `CaptureCache` | 3 |
Expand Down Expand Up @@ -67,7 +67,7 @@ The badge in the [README](../README.md) shows the automated test total. Harness
| `ResponseEconomy` | 18 |
| `RuntimeLaunch` | 11 |
| `RuntimeLogs` | 12 |
| `RuntimeOutput` | 5 |
| `RuntimeOutput` | 6 |
| `RuntimeRouting` | 65 |
| `RuntimeSessions` | 31 |
| `RuntimeTree` | 1 |
Expand Down Expand Up @@ -101,7 +101,7 @@ The badge in the [README](../README.md) shows the automated test total. Harness

## Python contract suites (`tests/test_*.py`)

**534 tests.** Derived from `test_*` methods on every `unittest.TestCase` subclass, read with `ast`.
**540 tests.** Derived from `test_*` methods on every `unittest.TestCase` subclass, read with `ast`.

| File | Tests |
| --- | ---: |
Expand All @@ -119,7 +119,7 @@ The badge in the [README](../README.md) shows the automated test total. Harness
| `test_elastic_ingress.py` | 13 |
| `test_elicitation_confirmation.py` | 6 |
| `test_engine_identity.py` | 5 |
| `test_field_trial.py` | 143 |
| `test_field_trial.py` | 149 |
| `test_follow_ups.py` | 5 |
| `test_fuzz_target_lists.py` | 6 |
| `test_initialize_instructions.py` | 5 |
Expand All @@ -142,15 +142,15 @@ The badge in the [README](../README.md) shows the automated test total. Harness

## Live Godot harness (`tests/*.ps1`)

**1386 assertions.** Derived from `Assert-True` call sites. The harness is one long live scenario rather than a set of named cases, so this counts assertions and says so.
**1388 assertions.** Derived from `Assert-True` call sites. The harness is one long live scenario rather than a set of named cases, so this counts assertions and says so.

| File | Assertions |
| --- | ---: |
| `bounded_reads.ps1` | 4 |
| `follow_ups.ps1` | 2 |
| `observed_post_state.ps1` | 12 |
| `refusal_remedies.ps1` | 3 |
| `run_godot_integration.ps1` | 1365 |
| `run_godot_integration.ps1` | 1367 |

## What is not counted here

Expand Down
2 changes: 2 additions & 0 deletions docs/TOOL_REFERENCE.md
Original file line number Diff line number Diff line change
Expand Up @@ -1134,6 +1134,8 @@ Paths are dot or slash separated, `architecture.inventory.slots` or `architectur
- `blackboard_read` (optional `board`, `path`, `deep`, `include_metadata`). `deep` returns the whole subtree; `deep: false` returns one level with nested containers replaced by a `_truncated` marker carrying their size, so a caller can see there is more rather than being handed a partial picture that looks complete. A read that finds nothing says which kind of nothing: `reason: "expired"` with `expired_at_ms`, and `expired_author` and `expired_reason` where the write supplied them; `reason: "cleared"` with `cleared_at_ms`, `cleared_by` and `cleared_reason`; or `reason: "no_record"`, which carries `last_board_clear` when the most recent board-level event was a clear of everything. A lapsed claim, a path somebody removed and a path with a typo in it lead three different ways and used to answer identically. A board remembers its 256 most recent expiries and 256 most recent clears.

Every read also returns `revision`, the board's, and `updated_at_ms` for a named path, which are what the two writers below pin a change to.

Every call that changes the board answers from the board it saved, read back from the file after the save: `revision`, a write's `metadata` and the `value` now stored at its path, and a task call's `task`. What an answer reports is what the next read sees (#1019).
- `blackboard_patch` (`operations`, optional `board`, `author`, `reason`, `expected_revision`). RFC 6902, applied all or nothing against the board root. If any operation fails, none is applied and the board is exactly as it was, and the refusal names the entry that stopped the batch: which index, which pointer, and for a failed `test` what the board actually holds. `path` and `from` are JSON pointers, so a board path like `doc.items` is written `/doc/items`.

**Writing against another agent.** The board exists because more than one client is expected, and until now the only concurrency guard on it was the task lease. Two agents that both read a key, both incremented and both wrote left the board holding one increment, with neither call an error. `blackboard_write` takes `expected_updated_at_ms`, the value a read reported for that path, or `0` for "and it must not exist yet"; `blackboard_patch` takes `expected_revision`, the board's, because a patch spans paths. A mismatch is refused `409` with `reason_code` `stale_write` or `stale_patch`, naming what the board actually holds and who last wrote it -- the shape `blackboard_task_claim` already uses for a claim somebody else is holding. A caller that passes neither keeps last-writer-wins.
Expand Down
36 changes: 36 additions & 0 deletions include/didi/common/terminal_text.hpp
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
#pragma once

#include <string>

namespace didi {

// Text with the terminal control sequences taken out: carriage returns, and
// escape sequences, CSI (colour, cursor movement) and two-character alike.
//
// Godot colours its console output, and Didi hands some of it to agents: an
// export failure used to arrive as four kilobytes of escapes and progress bars
// with the one actionable line sixty lines down (#651), and runtime_read_output
// relayed the colouring on 40 of a fresh headless editor's first 41 records
// (#1028). A client that renders either into a terminal would execute them,
// and a match against the visible text fails.
inline std::string withoutTerminalEscapes(const std::string& text) {
std::string out;
out.reserve(text.size());
for (size_t index = 0; index < text.size(); ++index) {
if (text[index] == '\r') continue;
if (text[index] != '\x1b') {
out += text[index];
continue;
}
// CSI and the two-character sequences alike: skip to the byte that ends
// the sequence rather than trying to understand it.
++index;
if (index < text.size() && text[index] == '[') {
++index;
while (index < text.size() && !(text[index] >= '@' && text[index] <= '~')) ++index;
}
}
return out;
}

} // namespace didi
5 changes: 4 additions & 1 deletion src/gdextension/gdextension_entry.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -176,7 +176,10 @@ GDE_EXPORT GDExtensionBool didi_library_init(GDExtensionInterfaceGetProcAddress
r_initialization->deinitialize = didi::godot::deinitialize_offline_helper;
r_initialization->userdata = nullptr;
r_initialization->minimum_initialization_level = GDEXTENSION_INITIALIZATION_CORE;
DIDI_LOG_INFO("GDEXTENSION", "Didi runtime disabled for isolated offline helper");
// DEBUG, because the helper's output is what shader_check_compile and
// its siblings hand back as raw_output, and at INFO every answer began
// with this line (#1028). The WARN default below is never reached here.
DIDI_LOG_DEBUG("GDEXTENSION", "Didi runtime disabled for isolated offline helper");
return 1;
}
// The extension's console lines go wherever the engine's stderr goes: the
Expand Down
Loading
Loading