Skip to content
OBSERVATION NOTES

Open-source Engineering

Runtime, workflow, recovery, tools, skills, observability and governance mechanisms.

99observation notes
Observation note list99 shown, sorted newest first
01
Daily ObservationExperiment ReportAwaiting review

The rules are written. Why might an agent still make a different choice?

An empty list of permitted directories returns allow. Three bounded experiments follow rule meaning through configuration, triggers and restored history.

→
02
Daily ObservationExperiment ReportAwaiting review

The secret came from an environment variable. Why was it still in the process list?

One synthetic token, three transport paths, and one process observer show why secret origin and secret transport are different claims.

→
03
Daily ObservationExperiment ReportAwaiting review

When an AI team fails, who checks the checker?

One of two answers is wrong, yet the score is 100%. Original evaluation code and public annotations show why tools that diagnose agent failures need checks of their own.

→
04
Daily Observationcomparative-researchAwaiting review

When Multi-Agent Systems Go Out of Control: What Agent Governance Do We Need in 2026?

Microsoft's runtime governance and a real OpenAI SDK fix lead to the questions of evidence, responsibility and authorization across agent handoffs—and why FCoP uses files for collaboration contracts.

→
05
Daily ObservationExperiment ReportAwaiting review

You approved the arguments. Did the tool execute the same ones?

We ran the same six cases against OpenAI Agents Python v0.22.2 and v0.22.3 to see how defaults, coercion and validators affect the approval boundary.

→
06
Daily ObservationExperiment ReportAwaiting review

More backups. Less history. How does that happen?

A full backup list can hide an empty history. Running the original merge code shows how unchanged saves can evict the version you actually need.

→
07
Daily ObservationExperiment ReportAwaiting review

No reply. Will retrying launch a second AI agent?

The action may have happened even when its reply disappears. A real-file experiment separates replaying a recorded result from refusing an uncertain retry.

→
08
Daily ObservationExperiment ReportAwaiting review

The quota is exhausted. Why is the agent still retrying?

A provider can name the reset time while a recovery system sees only a failed turn. We trace the same message through two versions of the classifier.

→
09
Daily ObservationExperiment ReportAwaiting review

Back in the same workspace, but which visit owns the result?

You leave workspace A, visit B, and return to A. A request from the first visit finally arrives. A controlled timing experiment asks whether it still belongs on screen.

→
10
Daily ObservationExperiment ReportAwaiting review

Resume the conversation. Keep the old credentials too?

A conversation needs continuity, but its launch credentials need not be permanent. A real-file experiment shows why resuming with a new value does not erase the old one on disk.

→