Skip to content
Experiment Report
Awaiting weekly review
Six AIs, One Task: What Happened?
RESEARCH · SIX AI TEAMS

Six AIs, One Task: What Happened?

Six AI teams tested through CodeFlowMu and FCoP: PM assignment, delivery quality, efficiency and recovery, checked against independent EVAL and retained file evidence.

Tested 2026-09-09File evidence · Independent EVAL · Six dimensions简体中文 →

The same inspection task. One AI team delivered in 12 minutes 20 seconds. Another was still unfinished after two hours. One never assigned a single child task.

Was the difference model capability, integration, or how the PM organized the work? We gave six AI teams the same assignment through CodeFlowMu and FCoP, then followed the retained task, execution, report, authorization, and independent EVAL records to find out.

6 modelsSame task · Multiple team roles
4 deliveries2 forced stops; failure records retained
12m 20sFastest normal delivery · Formal receipts

← All research articles · Public sources and attachments

Using CodeFlowMu and FCoP to make AI teamwork observable, documented, and open to scrutiny

Runs: September 9, 2026 · Analysis: September 10, 2026

On September 9, 2026, we used CodeFlowMu to give six AI teams the same system-inspection task. Each PM had to divide the work, use tools, collect specialist reports, and report to ADMIN. The central question was not simply which team finished fastest. It was what evidence CodeFlowMu captured, and how that evidence let us compare, diagnose, and judge the work. FCoP provided formal coordination records; CodeFlowMu connected tasks, execution, delivery, approvals, recovery, and independent EVAL analysis. Both successful and failed attempts remained available for review.

1. CodeFlowMu, FCoP, and the Test

1.1 What CodeFlowMu and FCoP do

CodeFlowMu is a multi-agent team coordination and governance system built on FCoP, with locally retained work records. It organizes agents in different jobs and supports formal delivery, evidence checking, approval, independent evaluation, runtime diagnosis, and evidence retention. Users can trace a task to its execution, then trace a report's claims back to supporting records.

FCoP, the File-based Coordination Protocol, expresses collaboration through formal files. TASK, REPORT, and related review records represent assignments, deliverables, and decisions in a persistent, inspectable form. CodeFlowMu applies that protocol to operating AI teams. See the public FCoP project.

Together, they address a practical problem: after someone asks an AI team to “check the system,” who actually checked it, what was examined, and what supports the claim that the work is complete? Task relationships, execution receipts, and reports make these questions answerable after the conversation has ended.

Who executes, and who evaluates?

The test configuration had one human ADMIN, four execution-team seats, and an independent EVAL seat. The table describes job responsibilities. Each PM decided how to divide the five inspection areas between DEV, OPS, and QA; those differences were part of the test.

RoleResponsibilityMain evidence
ADMINHuman requester and final acceptor; assigns work, authorizes restricted actions, and can terminate a runRoot submission, authorization, acceptance, and archive records
PM-01Interprets the request, creates assignments, coordinates progress, accepts child work, and reports to ADMINChild TASKs, dispatch and acceptance receipts, final PM REPORT
DEV-01Performs assigned technical, tool, or code-related inspectionsTool/execution records and DEV REPORT; inspection is not automatic authorization to modify code
OPS-01Checks assigned runtime, environment, and configuration concernsRuntime evidence and OPS REPORT
QA-01Performs assigned verification and review, distinguishing supported conclusions from issues and unknownsQA REPORT and its actual verdict
EVAL-01Independently examines the run, delivery quality, and evidence consistencyObservation report, run-record analysis, and evaluation-attempt records

All six runs used the same EVAL configuration: Cursor SDK / auto-smart. EVAL did not switch with the tested team. PM, DEV, OPS, and QA were the execution team; EVAL ran in its own session, did not perform the team's inspection assignments, and did not make PM or ADMIN acceptance decisions. Holding evaluator configuration constant reduced one source of variation. auto-smart is a routing label, however, not proof that the underlying foundation model stayed fixed.

EVAL has three distinct report paths: a task-run record analyzes a complete execution, a system observation examines system assets, and a closeout observation checks PM’s final report against its task evidence chain. All contain analysis, with different triggers, scopes and skill combinations; see the report and skill matrix in section 1.10. We check EVAL against original records and retain failures, retries and mismatched material. A common evaluator configuration does not mean that every run produced all three reports.

Original task panel: Doubao run

English reading guide — original Chinese UI: “待 ADMIN 验收” means “Awaiting ADMIN acceptance”; “子任务已完成” means “Child task completed”; “未投递 REPORT” means “Undelivered REPORT.” The screenshot preserves the conflict between completed tasks and stale queue labels.

Click the screenshot to inspect its original pixels.

Original screenshot, September 9: the root task awaits ADMIN acceptance while three specialist tasks are complete. The stale “undelivered report” message at the bottom was separately recorded as a display issue. A UI label cannot replace formal delivery receipts. Screenshots preserve the original Chinese interface; the surrounding English text explains the relevant evidence.

1.2 The exact task

All six runs received the same initial task body. Reminders, authorizations, and manual termination were recorded separately; this was not presented as an intervention-free experiment.

Original ADMIN task, preserved verbatim:

检查FCoP落地情况,检查MCP工具情况,检查SKILLS分配和使用,检查各角色权限和职责,检查轨机运行情况;请PM分解任务,团队协作完成;最后形成报告向ADMIN汇报!

English translation:

Check the implementation of FCoP, the MCP tools, the allocation and use of SKILLS, each role's permissions and responsibilities, and runtime operation. PM should decompose the task and have the team complete it collaboratively, then produce a report for ADMIN.

Task titles identified the tested model—for example, “system inspection codex,” “system inspection Doubao,” or “system inspection Qwen.” The shared test was the body above. The original wording “轨机” is retained in the source rather than silently edited.

Why this task?

First, it inspects the environment the team itself uses: coordination records, available tools, skill allocation, permissions, and runtime health. Second, it explicitly requires decomposition, teamwork, and a final report, exposing organization as well as individual reasoning. Third, it leaves realistic natural-language choices open. We did not preassign every checklist item or dependency, allowing us to observe how each PM chose scope, sequence, and a stopping point.

A reasonable deliverable is an evidence-backed inspection report. Each of the five areas needs a scope, a verification method, findings, and any unverified items. Workers should provide formal reports and the PM should form an overall judgment. Finding a system issue does not automatically mean the inspection failed. Conversely, “inspect and report” does not automatically authorize code repair. Inspection completion and product health are separate outcomes.

1.3 CodeFlowMu as the test instrument

CodeFlowMu organized the work and retained its evidence. It did not replace the PM's business judgment.

Test structure

The workflow was ADMIN → PM → DEV/OPS/QA → PM delivery, with independent EVAL examining the records. The instrument had three roles: operate the collaboration, collect evidence of what happened, and support independent review. All timing, completion, and quality comparisons follow that record chain.

1.4 Models and integration paths

“Six AIs” refers to six models operating as teams under their respective integration configurations. They ran in this order: Codex, Doubao, DeepSeek, Kimi, Qwen, Cursor. Model identifiers came from the identity replies and saved run records.

SchemeModel identifierProvider and execution framework
Codexgpt-5.6-terraChatGPT subscription / Codex CLI (H2 · Shadow)
Doubaodoubao-seed-2-0-pro-260215Ark API / Codex app-server
DeepSeekdeepseek-v4-proDeepSeek API / Codex app-server
Kimikimi-k3Moonshot API / Codex app-server
Qwenqwen3.8-maxDashScope API / Codex app-server
Cursorauto-smartCursor SDK

The first five used the Codex framework; Cursor used its SDK. The Cursor QA seat initially had a model unavailable through that Host, and recovered after authorization. That intervention remains part of the result. The EVAL configuration stayed Cursor SDK / auto-smart throughout.

Original model configuration page

English reading guide — original Chinese UI: the top section is “EVAL channel and model.” It shows the PM team using codex/gpt-5.6-terra, EVAL using cursor/auto-smart, and no active EVAL session. The bottom cards are PM (project manager), DEV (developer), OPS (operations), QA (quality assurance), and EVAL (independent evaluator).

Click the screenshot to inspect its original pixels.

This later-supplied configuration screenshot illustrates the product, rather than proving every historical run's configuration. It shows a Codex/gpt-5.6-terra PM team and a separately configured Cursor/auto-smart EVAL seat. No EVAL session is running in the screenshot. Historical model identity must still be checked against the corresponding run.

1.5 Initialization, execution, and timing

Every run began after CodeFlowMu system initialization. Records were exported and backed up before initialization for the next run. We configured and checked the next team's model, restarted the service on port 18766 after switching models, and submitted the same task. Late EVAL reports were preserved in supplementary backup batches.

ConditionProcedure
Machine and codeSame machine, CodeFlowMu V2.2.9, commit cb590ce
Starting pointSystem initialization before each new run
Execution teamSwitch and verify the tested configuration; record mistakes and recovery
EvaluationKeep EVAL at Cursor SDK / auto-smart
TaskSame original body in section 1.2
Scheduled reminderLeave the initialization-default PM progress reminder enabled; it wakes the PM but makes no business decision
RetentionSeparate export and backup by run; correlate sessions as well as possibly reused task IDs

Initialization is not a full snapshot restore of the OS, every cache, or external model services. Network and provider conditions can change. The result therefore compares integrated teams following a shared initialization procedure, not isolated foundation models under perfectly identical conditions.

Duration runs from the PM's first formal session to successful submission of its normal final report. Authorization and recovery within the run count toward elapsed time. Subsequent ADMIN acceptance and EVAL generation do not. For incomplete runs, we report time to forced termination, not “completion time.”

1.6 From a conversation to inspectable work

How CodeFlowMu supports the test

CodeFlowMu placed assignments, execution, delivery, and authorization into a traceable workflow. The PM's organizational choices became observable: how quickly it delegated, which relationships it created, whether reports were complete, and how it handled trouble. Similar final answers need not mean equally reliable work.

The Cursor run retained failed QA launches, ADMIN authorization, cancellation of the old task, the linked rerun, and final delivery. Qwen retained a different sequence: server-side task creation applied despite transport errors, and a complete OPS report body that failed to become a formal report. This distinguishes saying something is done, attempting to deliver it, and actually delivering it.

The same records also revealed instrument defects: stale projections, report-generation problems, and incorrect raw-material associations. That ability to investigate the instrument itself is valuable; it does not mean every projection is already correct.

1.7 “Files are the truth” means work must be written down

Our analysis uses a retained experimental archive, not an agent's memory or retrospective account. Original CodeFlowMu/Host execution records, FCoP artifacts, independent EVAL reports, and the backups, timelines, checks, and score tables built from them form the evidence base.

LayerRetained materialPurpose
Original runTASK, REPORT, sessions, tool returns, public progress text, approvals, issues, logsPreserve assignments, actions, claims, and authorization as observed
Independent evaluationBoth EVAL report types, failed generations, late supplementsSupply judgments and gaps that can themselves be checked
PreservationSeparate ZIPs, inventories, sizes, SHA256 hashes, source indexPreserve versions and avoid mixing runs
AnalysisStage timeline, timing tables, fact checks, diagnosis, scores, detailed reportTurn retained observations into reviewable conclusions

We did not ask agents to remember what they had done and then score the recollection. We reconstructed the run, checked task/report links, and compared statements with execution. These files remain useful after sessions end or models change.

Written records and evidence

The principle is durable documentation: assignments, actions, approvals, delivery, and analysis must be saved, transferable, and traceable. Files can still contain mistakes. A hash fixes the saved bytes; it does not establish the truth of a business claim.

Each run was exported as a separate raw-evidence.zip, with an inventory, sizes, hashes, and ZIP integrity checks. Late Qwen and Cursor EVAL outputs have separate supplements. Runtime evidence comes from CodeFlowMu and its Hosts; independent backup and later review are preservation and analytical steps applied to that evidence.

When records conflict, first align model, time window, Session, task, and report. Initialization can reuse TASK-001 or CUSTOM-001. A shared ID alone does not identify a run. Report text does not prove submission; a “running” display cannot overturn a cancellation receipt; a “pass” claim cannot substitute for the corresponding test.

1.8 The business support around agent work

Execution, verification, acceptance, EVAL, and diagnosis

CodeFlowMu supports what happens after an agent starts working: evidence checking, acceptance by responsible roles, independent observation, and diagnosis. These are connected by actual files. Execution status, verification, PM judgment, and EVAL analysis should not collapse into one vague “the system says complete.”

First, formal submissions and child tasks identify who is responsible. Execution records connect tasks, roles, sessions, calls, returns, and timestamps. Reports have formal writes and delivery records, separating execution, generated text, submission, and acceptance.

Second, REVIEW-GATE creates fact-check records. In the Cursor archive, a review of DEV task 002 includes execution_evidence_state: verified, review_state: needs_pm, business_decision: false, and attention_owner: PM. The evidence had been checked, but business judgment remained with the PM. third_party_source_state: not_configured does not pretend an external source was consulted. Compatibility fields must be read alongside the actual decision owner.

Third, PM accepts child work and ADMIN accepts the root task. Commands have receipts and state transitions; restricted actions have authorization records. EVAL and programmatic checks do not replace that responsibility. Cursor's recovery used this authorization path.

Fourth, EVAL records its observations and attempts. Diagnosis follows task, session, tool, delivery, and approval evidence to determine where a failure occurred. Confirmed issues can become ISSUE records; incomplete evidence remains a hypothesis rather than automatically becoming a product defect.

Support stageActual file or file family in the Cursor archiveWhat can be checked
Submission and assignmentsSUBMISSION-20260909-001.json, root/child TASKs, fcop/ledger/tasks.jsonlRequest, division of labor, cancellation and rerun relationships
Executionactions-20260909.jsonl, runtime-events-20260909.jsonlActions, sessions, results, and time
Tool transport.codeflowmu/logs/tool-transport-events.jsonlCall identity and transport events
DeliveryFour formal REPORTs, .codeflowmu/report-delivery/acks.jsonlContent and delivery to a PM session
Fact checkingREVIEW-…-REVIEW-GATE-on-TASK-….mdEvidence state, pending decisions, snapshots, responsibility
Decisions and authorizationTask-command receipts, approval audit, GOV recordRequests, applied operations, authorization scope
EVALOBSERVATION-20260909-002-panel-scan.md, …003-benchmark-CUSTOM-20260909-001.mdIndependent asset and run analysis
EVAL recoveryeval-observation-attempts.jsonlStarts, failures, retries, and outcomes
Issue trackingISSUE-20260909-001-PM.md and closure recordsHow the configuration issue was raised and handled
Process and usageChat/task JSONL, usage JSONLPublic progress, Host results, and usage context

The public evidence inventory identifies 26 verified files by name, size, and hash. The complete private logs and credentials are not distributed with this article.

1.9 What real artifacts look like

These are excerpts from the final Cursor archive, not invented templates. YAML fields and report sections are selected for explanation. Omitted fields may contain additional state and findings.

yaml
task_id: TASK-20260909-005
root_task_id: TASK-20260909-001
sender: PM
recipient: QA
parent: TASK-20260909-001
depends_on: []
acceptor: PM
rerun_of: TASK-20260909-003
subject: 只读核查 FCoP 落地与角色权限职责边界(模型修复后重跑)

recipient assigns QA, parent and root_task_id connect the root task, and acceptor assigns PM acceptance. rerun_of links new task 005 to old task 003. An empty depends_on indicates no explicit execution dependency. The excerpt establishes identity and relationships, not a passing inspection by itself.

yaml
kind: fact_check
task_id: TASK-20260909-002
report_id: REPORT-20260909-001-DEV-to-PM
review_state: needs_pm
execution_evidence_state: verified
business_decision: false
attention_owner: PM
third_party_source_state: not_configured

Verified execution evidence and a pending PM judgment coexist. The review explicitly says it made no business decision. This is how a saved fact check can support acceptance without silently replacing the accepting role.

markdown
## 子任务回执
- `REPORT-20260909-001-DEV-to-PM.md``TASK-20260909-002`(approved)
- `REPORT-20260909-002-OPS-to-PM.md``TASK-20260909-004`(approved)
- `REPORT-20260909-003-QA-to-PM.md``TASK-20260909-005`(approved;`rerun_of` 作废的 `TASK-20260909-003`

## 说明
全程保持只读巡检目标;PM 未修改业务代码/配置。根任务业务验收与归档由 ADMIN 决定。

The PM lists three worker reports and preserves the QA rerun and ADMIN acceptance boundary. The original report text is Chinese and is intentionally retained as evidence: it says DEV, OPS, and the new QA task were accepted, and final root acceptance and archiving belong to ADMIN. These are claims to check against receipts, approvals, and execution—not proof merely because the report says “approved.” See the excerpt source index.

1.10 EVAL: another agent observes, with three distinct report paths

When the execution team says “done,” who checks that claim? In CodeFlowMu, another EVAL agent reads the evidence. PM organizes delivery; DEV, OPS, and QA perform the work. EVAL uses a separate session to inspect their files and execution records, compare claims with evidence, and produce an independent analysis.

All six runs used Cursor SDK / auto-smart for evaluation. Independence means a separate role, session, and responsibility. Even when the tested team also used Cursor, EVAL was another session. This does not establish a different underlying foundation model or guarantee an unbiased, correct judgment. A common configuration helps comparison; evidence checking remains necessary.

Recording starts with the task; model analysis runs when a report is generated

When ADMIN enables recording at task creation, CodeFlowMu creates a run identity and retains the root task, environment and model baseline, start time, and starting log cursors. Existing chat, public progress, tool, session, and business logs retain the execution trail. The EVAL model is not necessarily running continuously or analyzing every event in real time.

Generating the task record freezes the end boundary, recursively identifies the complete child-task tree, gathers the run's material, and starts independent EVAL analysis. This separates continuous evidence retention from model-analysis calls. A run without recording enabled cannot later be presented as a complete prospective record. Existing raw records also do not prove that a final analysis report was generated successfully.

The report matrix: execution history, system assets, and delivery claims

ReportTrigger and scopeWhat its analysis answersTypical file marker
Task-run recordRecording enabled at creation; generation freezes one root task and its complete treeHow PM assigned work, what happened, timing, retries, interventions and outcomes; evidence and handling of linked ISSUEs*-benchmark-<RUN-ID>.md
System observationThe EVAL observation action and related observation triggers; a system-asset snapshotAsset consistency, risks, evidence gaps, emergence patterns, and candidates for knowledge retention*-panel-scan.md
Closeout observationThe root task's EVAL action; PM's final report and its evidence chainWhich delivery claims are supported, contradicted, or unverified, and what ADMIN should inspect before acceptance*-eval-agent.md

All three contain analysis. They examine different objects: how a run unfolded, what the system assets show, and whether a delivery report's claims have support. They are not renamed copies of one summary or three automatic approval gates.

System observation covers nine asset classes: ledger, runtime logs, public thinking/progress logs, usage, analytics, internal EVAL, emergence log, role views, and shared knowledge. Public progress means content actually emitted and retained; it does not expose or justify speculation about hidden reasoning.

Which skills support the reports?

Skills are work instructions for the EVAL agent, not additional models. One skill does not correspond to one report. The routing in test commit cb590ce combines seven skills. A checkmark means required injection for that path, not proof that a particular run performed the analysis correctly.

Skill and purposeTask recordSystem observationCloseout
eval-statistical-analysis: freeze identity; calculate timing, calls, failures and retries with formulas, denominators and provenance
eval-issue-analysis: examine linked ISSUEs, evidence status, impact, and confidence in causal hypotheses
eval-observation-writing: separate facts, inference, gaps and advice; cite sources and follow the output contract
eval-admin-closeout-observer: reconcile PM's final report, worker reports and the current authoritative evidence chain
controlled-emergence-observer: inspect task relationships for probes, self-tasks, sandboxes and project-tree patterns
eval-risk-gap-analysis: compare expected and actual behavior and recommend risk severity and ownership
eval-promotion-advice: recommend findings for follow-up tasks, issue drafts or reusable knowledge

eval-statistical-analysis governs calculation and interpretation: establish the task tree and time window, normalize and deduplicate events, then compute metrics. Separate active execution from waiting and assignment completion from product QA success. Missing evidence is unknown; conflicting evidence is disputed. Scoring is optional and requires an explicit formula; a high score cannot erase incomplete work or insufficient evidence.

The system-observation path also injects the closeout evidence-review skill, but its scope remains controlled by the system_observation contract. It does not thereby produce a separate closeout report. The repository also contains an auxiliary fcop-eval-promotion workflow for subsequent classification and internal drafts; it is not in these three paths' required skill lists. Advice does not automatically create tasks, publish issues, or change lifecycle.

The EVAL record matrix: logs, retained assets, and factual assessment

Read this system in three layers: logging retains what happened; asset organization makes tasks, reports, evidence bundles and indexes traceable; independent factual assessment checks claims, identifies contradictions, explains uncertainty and recommends follow-up. CodeFlowMu and the Host produce most raw logs; they are not all generated by EVAL. EVAL's analysis then becomes another retained file asset that can itself be reviewed and reused.

Factual assessment is a sourced, challengeable judgment, not an automatic declaration of truth or ADMIN acceptance. The seven skills make the matrix operational through statistics, issue analysis, writing, delivery verification, emergence observation, risk analysis and retention advice.

Evidence layerFile or fieldWhat it establishes
Recording start.codeflowmu/eval-recordings/<RUN-ID>.jsonRun identity, boundaries, task association and generation state
Run materialraw/, artifacts/ and indexes under research/evidence/benchmarks/.../<RUN-ID>/Material available to EVAL and later reviewers
Program collectionCollected run record, asset scan, *-evidence-bundle.mdWhat software gathered; not an independent agent conclusion
Independent analysisEVAL Session, model provenance, analysis_skill_ids or skill_ids, and skill receiptsWhich session received which instructions; injection proves loading, not analytical quality
Final reportAnalysis and agent/session/run provenance under fcop/internal/eval/Persisted independent analysis whose citations can be checked again

The EVAL agent reads evidence and returns analysis text. Runtime checks the required format, provenance and applicable skill receipts before persisting it. Program collection, agent analysis and final persistence are separate stages. A file appearing, a session ending, or a “generated” label alone is insufficient proof of a valid completed analysis. See the actual code excerpts in section 3.9.

CodeFlowMu provides both the execution evidence and a business workflow for another agent to question and analyze it. EVAL advises; PM remains responsible for delivery; ADMIN retains acceptance and follow-up decisions. This article additionally checks EVAL against the separate per-run backups. Three implemented report paths do not mean that every run successfully produced all three reports. Missing reports, failures, retries and mismatched material remain part of the evidence.

Implementation sources for this section are test commit cb590ce: codeflowmu-shell/src/eval-independent-analysis.ts for routing and required skills, eval-benchmark-recording.ts, packages/evaluator/eval-report-writer.js, and EvalObservationGenerator.ts. These explain the mechanism; source code alone does not prove a particular execution succeeded.

1.11 Why nine asset classes, and what does EVAL inspect?

The nine classes come from CodeFlowMu's implemented system-observation inventory. They let an evaluator compare how the same event appears in different records. They are not nine agents, nine scoring dimensions, or a universal taxonomy. The test version explicitly lists them in ASSETS_ANALYZED; business files such as TASK, REPORT and ISSUE enter verification through ledger associations, runtime material and report references.

AssetMain recordsAnalytical useQuestion to answer
1. LedgerRegistered tasks, role routes, parent relationships, states and business-record linksReconstruct formal work and compare lifecycle with deliveryWhat did PM assign, what was formally delivered, and what remains open?
2. Runtime logsSessions, attempts, leases, tool calls/results and errorsReconcile execution, retries, failures and state changesDoes a claimed action have a result? Did a finished session actually deliver a report?
3. Public thinking/progressEmitted plans, explanations and tool activityCompare contemporaneous claims with later behavior and scope changesDid promised dispatch occur? When did the plan change? Hidden reasoning is not inferred
4. UsageCollected request, token and usage recordsAttribute consumption where run associations exist and identify missing coverageWhich run incurred usage? Are retries or other sessions included? Without billing attribution, it is not exact task cost
5. AnalyticsAggregated events, counts and timing metricsCheck raw events, deduplication, denominators and windowsAre reported call counts and durations calculated consistently?
6. Internal EVALPrevious observations, collection material, analysis and generation statesCheck evaluation identity, provenance, persistence and mismatched materialIs this completed agent analysis or only a collected summary? Does it belong to another run?
7. Emergence logPreviously observed patterns, sources, risks and recommendationsCompare task relationships and track duplication, evolution or unsupported findingsWas this pattern already recorded? What changed, and is it reusable?
8. Role viewsRole-specific task lists and status projectionsCompare views with ledger, lifecycle and the view contractDoes “running” match formal state? Is an empty task list correct at this stage?
9. Shared knowledgeShared rules, experience, knowledge and reusable materialCheck relevance, age and opportunities for retentionWas available knowledge unused, or is useful knowledge missing?

The value lies in cross-asset verification. A PM claim of tool success requires a result, not only a call-start log. A “running” display that conflicts with cancellation receipts and session state needs projection checks. Runtime, usage and analytics help distinguish requests, retries and duplicate event records.

The scan retains four observation fields per asset: status, key finding, risk/value, and evidence. Independent EVAL then analyzes contradictions, gaps and recommendations. “All nine classes scanned” establishes inventory coverage, not that all nine were verified healthy. CodeFlowMu makes logs useful analytical assets by enabling cross-checks rather than merely accumulating files.

1.12 Emergence observation: inspect collaboration patterns, not only errors

EVAL also looks for task origins and organizational structures worth recording. “Emergence” has a bounded implementation here: it is not a claim that a model suddenly acquired new intelligence, nor a label for ordinary errors.

ObservationIdentification evidenceAnalytical purpose
Controlled emergenceRole routing combined with explicit probe, self-task and sandbox bootstrap markersEstablish where test/probe tasks came from and whether they contaminate business task lists, review queues or counts; retain probe evidence without treating it as delivery
Project-tree emergenceActual parent relationships and role routes within a thread, forming root → phase → executionObserve project/phase organization, assess value against outcomes, and check wrong parents, cross-thread links, cycles or closed parents with open children
text
Ordinary delegation: ADMIN root → PM assigns DEV / OPS / QA checks
Project-tree pattern: ADMIN main task → phase task → specialist execution tasks

A title containing “Phase,” “project” or “probe,” or simply creating more children, does not establish either pattern. Detection examines origin, parent, thread_key, routing and relevant markers. Independent EVAL checks the interpretation and reports alternatives, confidence, risks and advice. controlled-emergence-observer is one of the seven skills; the companion project-tree-observer.js supplies structural evidence, not another mandatory report.

Emergence observation examines both risk and value. Probe tasks appearing in business views may mislead users; a useful phase structure may provide a reusable way to organize work. Structure alone does not prove successful delivery. The system observation report and retained emergence log support later review and knowledge retention. EVAL recommends action; it does not autonomously clean up tasks, archive them or change dispatch.

No detected emergence is also a result worth retaining. Normal ADMIN→PM→DEV/OPS/QA delegation in this inspection cannot be marketed as new emergence merely because the team collaborated. Any per-run claim requires that run's report and relationship evidence. CodeFlowMu can accumulate observations about how collaboration structures form and whether they help, alongside success and failure records.

Implementation sources are the test version's packages/evaluator/eval-report-writer.js, controlled-emergence-observer.js, and project-tree-observer.js. The asset classes are an implemented inventory; the questions above explain analytical uses and do not assert that every question was verified in every run.

2. Overall Results and Each Team

The results below were checked against formal assignments, reports, execution, and cancellation records captured through CodeFlowMu. We separate completion, inspection quality, and integration failure.

2.1 The six results

SchemeNormal deliveryValid worker reportsElapsed timeScore / 100Main observation
CodexYesThree roles12m 20s88Fast closure with qualified conclusions
CursorYesThree roles, plus one cancelled old QA task20m 16s87Authorized recovery followed by delivery
DeepSeekYesThree roles39m 21s75Deeper collection, some incorrect interpretation
DoubaoYesThree roles21m 59s66Smooth workflow, overconfident report claims
QwenNoDEV only; PM submitted a blocked reportTerminated after 122m 10s40Useful diagnosis, ineffective closure
KimiNoNo downstream rolesTerminated after 31m 43s18Integration/session failure before teamwork

Scores for the six integrated runs

Weights: completion 25, result quality 30, efficiency 15, scope control 15, recovery 10, observability 5. Scores describe these runs, not universal foundation-model ability. Codex and Cursor's one-point difference places them in the same leading group. See the score data.

2.2 Completion is more than a final paragraph

Four teams delivered three valid specialist reports and a normal PM final report. Qwen's OPS content remained in a failed submission, QA produced no formal report, and Kimi created no child tasks. Forced archiving stopped a test; it did not complete the assignment.

Codex took fewer actions to reach a qualified conclusion. Cursor showed authorized recovery. DeepSeek's advantage over Doubao was mainly evidence and interpretation quality, not a different final task-state label. Qwen's 40 acknowledges diagnostic work as well as failed delivery; Kimi's 18 describes a non-delivering configuration. Neither score can be read as an intrinsic model score. This was one run per model configuration, not a blind or randomized benchmark.

2.3 Codex: fast delivery with explicit limits

The PM assigned FCoP/MCP to DEV, runtime to OPS, and skills/permissions to QA, without explicit dependencies. The first child task appeared around 1m 21s after the formal session began. Three worker reports and the PM summary covered all five requested areas.

Its strength was distinguishing static configuration from demonstrated capability and preserving unverified limitations. QA did not turn workflow completion into an unconditional product pass. The 12m 20s duration coexisted with 24/30 result quality. Some verification remained static, so this was not comprehensive functional certification.

2.4 Doubao: a smooth chain with weak final review

At about 1m 01s, PM began assigning DEV to MCP, OPS to runtime and permissions, and QA to FCoP/SKILLS. The three tasks had no explicit dependencies. Formal delivery took 21m 59s.

The quality deductions concern unsupported claims: 23 tools claimed but 22 listed, inconsistent skill-count scopes, a greater-than-95% match rate without a denominator, and treating visible schemas as proof of successful parameter validation. PM did not sufficiently filter worker overstatement. Completion 25/25 and quality 12/30 describe different aspects of this same run.

2.5 DeepSeek: deeper evidence, imperfect interpretation

PM created three role tasks without explicit dependencies, with the first at roughly 6m 29s and delivery after 39m 21s. Reports distinguished 48 skill references from 54 on-disk directories and on-demand injection. They included software probing and runtime endpoint checks.

However, the team misinterpreted the independently configured Cursor EVAL as a governance violation and recommended committing runtime governance files to the main repository. PM did not correct those interpretations. Quality 18/30 reflects stronger fact collection but insufficient interpretive review. Agreement among several roles does not make a claim true.

2.6 Kimi: the team never reached formal downstream work

Runtime recorded eight formal PM sessions ending as failed: seven error endings and a final cancellation. The first formal session received an encrypted_content content-type rejection. Six subsequent failure results contained overload and stream-disconnection messages. Tool-preparation records also showed readiness had not been achieved.

PM read resources and requested recovery, but created no formal child tasks or team reports. ADMIN stopped the run after 31m 43s. This establishes that the tested model/Host/adapter combination failed to deliver; it does not establish that Kimi intrinsically cannot use MCP or can never work through Codex. Its teamwork was not adequately exercised.

2.7 Qwen: investigation expanded without complete delivery

PM took 11m 46s to create the first child task, then assigned OPS, DEV, and QA, with QA referencing DEV. DEV delivered. OPS generated a complete report body, but its formal submission failed in transport. QA produced no formal report. PM submitted a blocked report and the run was terminated after 122m 10s.

The team encountered genuine transport and dispatch problems. Saying it found nothing would be inaccurate. Yet later probing, waking, pausing, and explanations did not produce effective resubmission or closure. Sparse early progress text and extensive later explanation also made timely supervision difficult. The 40-point result recognizes both diagnostic contributions and failure to deliver.

2.8 Cursor: authorized recovery left an inspectable history

DEV covered skills and development-side tools, QA covered FCoP and role permissions, and OPS covered runtime/MCP. QA's unavailable model caused four launch failures. PM requested recovery authorization; ADMIN approved. The old QA task was cancelled and new task 005 ran with a link to the old task. Once all three valid reports were present, PM completed delivery in 20m 16s.

Original Cursor task and rerun records

English reading guide — original Chinese UI: the root awaits ADMIN acceptance. Specialist rows show DEV, the original QA task, OPS, and a replacement QA task. The last task title includes “rerun after model recovery”; its existence alone does not prove the original QA succeeded. Cancellation and rerun receipts establish that distinction.

Click the screenshot to inspect its original pixels.

Four visible child-task rows did not mean four successful worker deliveries. Old QA task 003 was cancelled; task 005 supplied the valid QA report. A file in a done lifecycle directory must still be interpreted using its actual decision.

This was not fully unattended recovery: human approval was required. The authorization, cancellation, and rerun records are precisely what make the recovery explainable. The score of 87 reflects near-Codex delivery quality and organization under a real configuration problem.

2.9 How the score should be read

Scoring dimensions

If we judged only a done label, four teams would tie. If we judged only duration, unsupported conclusions would disappear. Completion, quality, efficiency, scope, recovery, and observability must be considered together.

Doubao lost quality points for unsupported proportions and overclaiming. DeepSeek lost them for incorrect governance interpretation accepted by PM. Qwen's completion 10/25 and efficiency 2/15 acknowledge partial delivery while recognizing the missing OPS/QA reports. Kimi's completion 0/25 records the absence of formal delivery, not a measured inability to write a good report.

Removing efficiency and normalizing the remaining 85 points gives Cursor about 88.2 and Codex about 87.1. Their order reverses but their grouping does not. A one-point difference in a single-run, judgment-based rubric is not statistically meaningful superiority.

The transparency concern was inconsistent public communication: Qwen had long early stretches of tool activity with little explanation, then much more text when stuck. Later explanation cannot restore the opportunity to supervise and stop earlier. The evidence does not establish an intention to hide activity.

3. What the Evidence Lets Us Analyze

This section traces PM assignments through TASKs, time through sessions and returns, quality through worker and PM reports, and EVAL conclusions back to original material. CodeFlowMu's linked records make “what differed, and why?” an inspectable question.

3.1 Compare the PM's entire delivery chain

The same task did not produce the same task graph. PM chose who worked first, who reviewed whom, and which relationships counted as dependencies.

PMAssignment patternActual deliveryPM review/summary assessment
CodexDEV: FCoP/MCP; OPS: runtime; QA: skills/permissions; no explicit dependenciesThree role reports and PM summaryQualified conclusions; static checks not presented as full functional proof
DoubaoDEV: MCP; OPS: runtime/permissions; QA: FCoP/SKILLS; no explicit dependenciesComplete team chainDid not adequately correct counts, scope, proportions, or validation claims
DeepSeekThree-role inspection without explicit dependenciesTeam reports with useful runtime collectionAccepted incorrect interpretations of EVAL configuration and runtime files
KimiNo formal child assignment; tool discovery, recovery, and session failuresNo team deliveryNo final report to assess; organizational capability not sufficiently exercised
QwenOPS inspection; DEV code-level diagnosis; QA verification referencing DEVDEV report, failed OPS submission, missing QA report, blocked PM reportSome diagnoses were corrected later; recovery did not yield closure
CursorThree specialist tasks, followed by cancellation and linked QA rerunThree valid reports plus retained cancelled taskAuthorized recovery is evidenced; human intervention remains part of the result

PM report quality means checking coverage of all five areas, evidence for important claims, correction of worker contradictions, explicit verified/unverified/failure distinctions, and recommendations within the mandate. Concatenating worker reports is not enough.

Kimi's missing final report is “not assessable,” not “a badly written final report.” Qwen's existing blocked report should be judged on obstacle explanation and proposed recovery. The article's result-quality score is a whole-run dimension, not a separately calibrated PM-writing score.

Assignment structures

Three independent tasks were not necessarily superficial. Codex and Doubao divided the same scope differently; either can be reasonable when evidence and coverage are clear. Qwen chose a heavier diagnostic route, beginning delegation later and making QA reference DEV. Time before the first assignment is itself an organizational choice.

Reference semantics matter. A reference should supply context without automatically blocking work. This run exposed a dispatch gate that treated that relationship inconsistently. That is a system issue. PM still has to judge whether the added relationship is needed and how to adapt when it obstructs a bounded inspection.

3.2 Visible configuration is not demonstrated capability

Levels of tool evidence

Resource visibility, schema visibility, actual tool execution, and successful business delivery are different levels of evidence. One cannot substitute for the next.

Codex retained limitations and a partial QA conclusion. Doubao treated schema listing as validation and mixed skill scopes. DeepSeek made useful distinctions in its collection but then misread the independent EVAL setup. PM review must test the relationship between each claim and its source, not reward file volume.

Kimi never reached a comparable complete team chain. Its actual resource reads and recovery calls matter, but so do the request rejection and interrupted sessions. The records support failure of this configuration, not a universal inability to understand MCP.

3.3 Did longer runs buy more useful evidence?

Elapsed time to delivery or termination

Codex's fastest delivery still had materials corresponding to all five areas. Cursor's 20m 16s included authorized recovery. DeepSeek's additional checks had value, but its report also contained a wrong governance interpretation. Qwen's two hours combined real infrastructure obstacles with ineffective later investigation.

The four completed runs had a median of about 21m 08s. A 20–30 minute budget is a reasonable starting hypothesis for a future bounded inspection on this environment. It was not a limit imposed on these runs, so it cannot be retroactively treated as a contract violation.

3.4 Efficiency is evidence value per unit of effort

Time versus result-quality score

Among completed runs, Codex's speed did not accompany a lower quality score. DeepSeek spent about 27 extra minutes and collected additional runtime evidence, but did not produce a more trustworthy overall conclusion.

Elapsed time includes preparation, pre-delegation checking, the longest worker path, retries, PM review, and summary. Stages overlap; role durations cannot simply be added. Cursor also includes authorization waiting. The evidence does not support a precise allocation such as “X% model delay, Y% platform delay.”

More issue claims must first be screened for false positives. Unsupported proportions, old snapshots, and incorrect governance explanations create review work. Issue counts and report length are not substitutes for useful findings.

Cost comparisons also need boundaries. Qwen's roughly CNY75.56 came from a daily screenshot. DeepSeek's CNY15.6486 export included other requests in its hourly aggregation. Subscription channels do not have zero cost. This evidence supports a time-efficiency assessment, but not a precise cost-per-qualified-report ranking.

3.5 Where Qwen's two hours went

Qwen timeline

QA was created at 17:09:51 and started at 18:02:15: 52m 24s waiting. The declared relationship was informational_reference, yet a dispatch gate required successful DEV closure. The historical source and runtime trace support this issue.

After QA started, it ran for nearly another 50 minutes without formal delivery. The earlier dependency no longer explains that entire delay. OPS called write_report with complete content, but received Transport closed; no formal report was saved. PM then probed, woke, and paused repeatedly without an effective resubmission path, ultimately reporting blocked.

Some deeper claims should be withdrawn. The alleged “ghost running QA” used a topology snapshot from before QA started. The allegedly random limit failures compared string and integer arguments. The root cause of short-name instability was not demonstrated under equivalent conditions. Genuine issues, hypotheses, and mistaken interpretations must remain separate.

3.6 Cursor's recovery was a governed sequence

Cursor recovery sequence

System initialization had occurred, but the QA seat was misconfigured with qwen3.8-max, unavailable through Cursor SDK. OPS identified a model-availability problem rather than an MCP or lease failure. PM requested authorization at 21:51:56; ADMIN approved around 21:55.

The old QA task was cancelled. A new task linked by rerun_of started at 22:00:12 and produced its report at 22:01:46. PM's successful final submission followed at 22:04:27. These relationships preserve the failed attempt rather than overwriting it as success.

Recovery of an authorized execution prerequisite is different from unauthorized scope expansion. In the September 8 comparison, the first Codex run expanded inspection into code repair and was stopped. That was a separate historical run. The September 9 Codex run did not repeat the behavior; this does not prove the earlier event was caused by another AI's old records.

3.7 Does the Codex integration explain failure?

Both models with incomplete runs used the Codex framework. Integration is therefore part of the causal analysis. The evaluated object is the model, provider API, adapter, Host, and CodeFlowMu working together—not a foundation model in isolation.

SchemeEvidenceSupported conclusion
KimiOne explicit rejection of encrypted_content; six later disconnections carrying an overload message; no downstream tasksA content-compatibility error and runtime failures occurred. Team capability was not sufficiently exercised to infer weak intrinsic ability
QwenCreated tasks, used tools, delivered DEV work; QA dispatch waiting, OPS submission failure, old-snapshot misinterpretation, extended investigation without closureIntegration could operate, but the delivery chain had faults and PM handling was insufficient. Compatibility does not explain everything

Kimi's overload messages are not themselves compatibility errors. Qwen's dispatch gate issue belongs to this system workflow and is not automatically attributable to Codex or Qwen.

Codex, Doubao, and DeepSeek completed through the same overall framework, so “uses Codex” is not sufficient to explain failure. Provider-specific request content, tool representation, returned results, and recovery behavior matter. Cursor's successful SDK run changed both model and framework, so it is not a controlled demonstration that moving Kimi or Qwen to another Host would fix them.

Integration paths

The tested source configured custom providers for Responses and disabled Codex reasoning-summary metadata support. The Doubao Ark bridge also normalized request fields and changed tool exposure. That Ark-specific behavior should not be attributed to every provider. Identical framework names do not guarantee identical effective tool catalogs, requests, or public output.

Disabling a summary does not establish that a model stopped reasoning or could not provide public progress. Qwen's sparse early explanation and later verbosity require comparison of raw assistant output with Host and UI events. The observed inconsistency harms supervision; deliberate concealment is not established.

To separate model ability from integration effects, repeat the same model and task on another verified executor with matched tool semantics and budgets. A website chat is not an equivalent control. These six runs do not quantify an integration penalty.

3.8 The evaluator also needs checking

Evaluation material identity

A uniform evaluator cannot compensate for incorrect inputs. Some task-record analyses described the current run but linked to raw packages containing morning Codex sessions, role routes, and report hashes. Initialization reused CUSTOM identifiers, creating a risk of mixing windows. coverage=complete did not establish identity correctness. This article used separate run backups rather than copying disputed counts.

There was also a report-generation problem: replay showed section extraction treating child headings as the end of a parent section, making an otherwise populated report appear empty. The UI sometimes displayed completed as a failure reason. A later successful Cursor generation proved that one output was accepted, not that the defect had been fixed.

Billing has an analogous identity problem. DeepSeek's export contained 168 requests, about 16.22 million tokens, and CNY15.6486, with hourly aggregation including identity chat. Qwen's daily screenshot showed about CNY75.56 without task-isolated detail. Unknown cost must not be replaced with zero or an invented precise ranking.

3.9 EVAL record reports and observation reports

EVAL used Cursor throughout. The following comparison focuses on system observations and task-run records; a separate closeout path checks PM’s final report and its evidence chain. The three paths are distinguished in section 1.10. Programmatic collection and independent analysis are separate stages. A generated collection record is not automatically a finished EVAL judgment.

01 / Collect the run material

The collector receives the project root and run ID. Collecting material is distinct from completing independent analysis.

typescript
  const child = spawn(
    process.execPath,
    ["packages/evaluator/eval-benchmark-record.cjs", "--project-root", projectRoot, "--run-id", runId],
    { cwd: projectRoot, detached: true, stdio: "ignore" },
  );

Source excerpt: codeflowmu-shell/src/eval-benchmark-recording.ts, lines 592–596, test commit cb590ce35686cb1980e3c89a7d68bd0cfbeb825a.

02 / Start a separate EVAL session

The target report, analysis kind and root task are bound to an EVAL session. The tested EVAL configuration was Cursor / auto-smart in every run.

typescript
      const handle = await runtime.sessionManager.startSession(
        agentId,
        sessionTaskId,
        {
          text: skillInjection.prompt,
          maxToolRounds: DEFAULT_SESSION_MAX_TOOL_ROUNDS,
          uiLang: readPanelUiLang(getProjectRoot()),
          context: {
            eval_observation: true,
            eval_analysis: true,
            analysis_kind: request.analysisKind,
            analysis_target_path: request.reportPath,
            root_task_id: request.mainTaskId,
            session_kind: "CHAT_BOUND",
          },
        },
      );

Source excerpt: codeflowmu-shell/src/web-panel.ts, lines 17863–17879, test commit cb590ce35686cb1980e3c89a7d68bd0cfbeb825a.

03 / Validate the analysis before accepting it

A completed session does not automatically constitute a valid report. This excerpt rejects malformed analysis; the surrounding function also checks provenance and required skill receipts.

typescript
  const contentErrors = validateEvalAssistantText(content, input.analysisKind);
  if (contentErrors.length) {
    throw new Error(`EVAL_ANALYSIS_FORMAT_INVALID: ${contentErrors.join(",")}`);
  }

Source excerpt: codeflowmu-shell/src/eval-independent-analysis.ts, lines 286–289, test commit cb590ce35686cb1980e3c89a7d68bd0cfbeb825a.

04 / Persist the report as UTF-8

After analysis and provenance are assembled, a temporary UTF-8 file replaces the target report. This code explains the mechanism; it does not prove that any particular run passed the checks.

typescript
  const temporary = `${absolute}.analysis-${process.pid}-${Date.now()}.tmp`;
  writeFileSync(temporary, raw, "utf8");
  renameSync(temporary, absolute);
  return absolute;

Source excerpt: codeflowmu-shell/src/eval-independent-analysis.ts, lines 393–396, test commit cb590ce35686cb1980e3c89a7d68bd0cfbeb825a.

EVAL findings across the six runs

The risk labels below come from independent analysis sections. High risk can concern evidence identity rather than poor team delivery; absent reports do not mean low risk.

SchemePanel observationTask-record analysisInterpretation
CodexLowMediumCompleted inspection with remaining evidence gaps; partial observation, not unconditional product pass
DoubaoMediumHighCurrent delivery existed, but frozen materials were disputed
DeepSeekLowHighTask/report inventory broadly consistent; raw package identity disputed
KimiFinal report not obtainedFinal report not obtainedUnderlying recording and execution evidence still exist
QwenMediumHighBusiness blocked and frozen-material identity disputed
CursorMediumHighLive recovery traceable, but linked task-record materials disputed

Doubao and Cursor EVAL analyses corrected a scanner's false alarm: after the root task moved to ADMIN acceptance, an empty PM todo view could be correct. The completed run must not be labelled missing solely from that projection.

Kimi had records even without a completed final EVAL pair. Its archive contains a 17,488-byte .codeflowmu/eval-recordings/CUSTOM-20260909-001.json, begun at 16:08:52 Beijing time, with state: recording and generation_attempts: 0 in that snapshot. Task/chat process files and Runtime events are also preserved. This establishes a recording entry, not successful final report generation. We did not fill the gap using another run's reports.

Original EVAL report generation screen

English reading guide — original Chinese UI: “生成任务记录” means “Generate task-run record”; “已生成” means “Generated”; “报告已生成” identifies the saved report; “查看记录报告” opens it. The background lists the standard task-run record and EVAL panel scan.

Click the screenshot to inspect its original pixels.

The screenshot shows a successful later Cursor report. Saving a report is one check; matching its content and identity to the run is another.

3.10 EVAL is useful, but not the final truth

Performance scores and EVAL risk are different measures

The article asks both whether the team delivered and whether its evidence is reliable. EVAL is valuable because it can challenge a convincing report. Its own claims also need review.

Cursor's score of 87 and high task-record risk are not contradictory measures of the same thing. The score uses the corresponding independent run backup and recovery chain. The risk addresses the associated record package's identity problem. We did not use disputed raw counts as if they belonged to the current run.

EVAL's corrections were often more useful than adding issues: it rejected “empty PM todo means lost work,” and distinguished actual Qwen transport failure from an asserted permanent dependency deadlock. DeepSeek's claim that the independent Cursor EVAL was a violation also needed correction: that configuration was intentional.

CodeFlowMu's product value is that successful delivery, authorized recovery, disagreement, and defects can all be investigated from retained records. Reliable run identity, correct freezing, and truthful generation states are priorities for improvement.

3.11 Fact checking: return a claim to its source

ClaimEvidence comparedFinding
“OPS finished the report”Complete call body, returned status, formal file and receiptQwen generated the body but failed submission
“PM todo is empty, so work was lost”Root state, ADMIN acceptance view, projection contractEVAL corrected the scanner's inference
“QA is running but has no session”Snapshot time, QA launch time, session recordA pre-launch snapshot cannot prove ghost execution
“This raw package belongs to this run”Model, time window, Session, routes, hashesSeveral same-named packages contained earlier Codex material

Identify the same run and attempt before asking what a record proves. Successful transport, successful command execution, and successful business delivery are different outcomes. Missing evidence remains unknown; conflicting evidence remains disputed until resolved.

No usable external third-party fact source was configured in this test. These checks relied on local execution evidence. They are not an external database certification or an automatic endorsement from a standards authority.

3.12 Diagnosis: make “stuck” a specific problem

Diagnosis should explain what happened, where the evidence is, and which layer needs investigation next. CodeFlowMu linked tasks, sessions, tool results, and logs to formal work, rather than leaving diagnosis as a conversational impression.

LayerObserved factDiagnostic implication
ConfigurationCursor QA used an unavailable modelAuthorized correction, cancellation, and a linked rerun restored delivery
SubmissionQwen OPS returned Transport closed after sending complete contentConfirmed submission failure; inspect recoverable delivery rather than blame only the model
DispatchQwen QA's reference relationship still waited for DEV at one gateExplains initial waiting, not all later non-delivery
Integration/sessionKimi content rejection and interrupted sessionsValidate the request and endpoint path before judging full team ability
EVAL outputMaterial collected but extraction, validation, or UI status problematicSeparate collection, analysis, and final file write

Useful diagnosis narrows the problem and supports recovery. Cursor supplied an end-to-end recovery example. Qwen supplied a case where actual issues were found but diagnostic scope was not controlled. Independent evaluation raises questions, fact checking tests their basis, and diagnosis identifies causes and possible next steps. PM and ADMIN retain responsibility for decisions.

3.13 The two incomplete runs, with precise evidence

Kimi did not establish downstream teamwork. Qwen established it but failed to finish delivery. ADMIN termination was the ending action, not the sole explanation for what preceded it.

For Kimi, the archived Runtime file contains an explicit error at 16:15:10: invalid_request_error: responses: unknown content part type: "encrypted_content". Six later failure results, from 16:19:18 to 16:33:12, contain responseStreamDisconnected and The engine is currently overloaded, please try again later.

The eight formal failed endings consist of seven error endings and one final cancellation. Four recovery calls returned refresh_queued with tools_ready:false. Resources and commands were used, but no formal child TASK or REPORT resulted. In the corresponding archived sdk.result records for Codex, Doubao, DeepSeek, and Qwen, these two specific signatures were not found. That comparison does not imply their tool paths were fault-free.

The content rejection is direct compatibility evidence. The overload text is a returned service signal, not an independent measurement of actual server load. We still cannot identify which component introduced or retained the unsupported content, the precise origin of the overload message, or whether the tool-readiness issue shares the same cause. See the timestamped source evidence.

For Qwen, the reference dependency's gate explains 52m 24s of waiting, and OPS transport failure explains a real delivery obstacle. QA then ran nearly 50 minutes without a formal report. Continued investigation included old-snapshot and argument-comparison mistakes. These support criticism of PM recovery and scope control, but do not quantify the share of blame attributable to each component.

Future validation should differ: Kimi first needs a minimal formal tool-discovery, child-task, and report chain; Qwen needs validated dispatch/submission recovery followed by a bounded inspection. These are proposed checks, not fixes already completed.

4. Conclusions: What CodeFlowMu Made Possible

4.1 The system made the comparison evidence-based

The demonstrated value is turning a natural-language request into teamwork that can be assigned, tracked, checked, and handed over. CodeFlowMu supplied operational visibility and management controls. FCoP made assignments and deliverables persistent. EVAL provided another opportunity to challenge conclusions.

Retained system evidenceQuestion it answersResult in this test
TASKs and relationshipsHow did PM divide the work?Distinguished parallel inspection, reference dependencies, and linked reruns
Sessions, returns, timestampsWhat actually ran, and where did it fail?Identified Kimi's content rejection and overload/disconnection sequence
REPORTs and receiptsWas content actually delivered?Avoided counting Qwen OPS text as formal delivery
Worker reports, PM summary, reviewsDid PM correct unsupported claims?Exposed Doubao's count/proportion/validation problems
Approvals, cancellation, rerunsHow was recovery authorized and executed?Reconstructed Cursor's recovery without erasing the failed QA attempt
EVAL records, reports, statusWhat was independently assessed, and what was missing?Preserved disputes and allowed Kimi diagnosis from underlying records

These judgments depend on evidence captured and retained during CodeFlowMu operation. Host/tool errors are recorded as returned; FCoP artifacts express formal collaboration; CodeFlowMu connects them to the business workflow. The exported backups and subsequent checks then turn records into analysis. This is not a story reconstructed from agent memory or from a green status label.

System defects remain in the article because the same records make them inspectable. Success has delivery evidence; failure has diagnostic traces; recovery has an authorization history; evaluation has sources. Scores and diagnoses are EVAL and analytical judgments, not automatic business decisions made by the system.

4.2 Overall assessment of the six models

For this inspection, Codex is the first choice for routine execution; Cursor belongs in the same group where authorized intervention is available. Among the domestic-provider models, DeepSeek is the first candidate for further testing, followed by Doubao. Qwen needs bounded task-control testing; Kimi needs integration validation first.

Codex and Cursor both scored 24/30 for result quality. Their advantage was keeping conclusions proportionate to evidence while reaching delivery. DeepSeek and Doubao showed they could organize the workflow, but PM review did not consistently filter incorrect interpretation or overstatement. Qwen and Kimi require different diagnoses rather than a shared “bad model” label.

4.3 Capability, reliable delivery, and review are separate thresholds

Doubao and DeepSeek demonstrated task decomposition, tool use, and formal team reporting through this framework. That does not establish that every domestic model is ready for an unsupervised PM role, nor does the failed pair prove that domestic models lack the underlying ability. Running successfully, reporting truthfully, and recovering effectively are distinct thresholds.

CodeFlowMu improvements should prioritize reference-dependency semantics, recoverable report submission, cross-run identity, and EVAL format handling, followed by stale display and progress-text issues. Traceability has been demonstrated; perfect calibration has not.

The original “inspect and report” wording did not authorize code repair. Production requests can further specify read-only scope, permitted formal assignments/reports, a budget, and when to stop investigating. Clearer language reduces ambiguity but does not replace PM judgment. A separate bounded-text test should not be mixed with the original wording and presented as an unexplained model improvement.

The scores remain 88, 87, 75, 66, 40, and 18 for these integrated runs. Foundation-model ability and integration loss were not separately measured.

4.4 A reliable team also knows when to stop

Inspection does not need to eliminate every defect. Reliable findings, supporting evidence, and sensible next steps can complete the assignment; repair needs its own authorization.

Proposed repeat-test procedure

A future A/B protocol could retain the original task in one arm and explicitly bound read-only work and time in the other. Rotate the order, repeat each model configuration at least three times, and preserve task records, system observations, any triggered closeout observations, and failed drafts. Exporting/checking old evidence and verifying the initialized next environment solve different problems.

A common Git commit is not a full machine snapshot. Provider, adapter, effective tool catalog, network, task graph, and interventions still differ. Verify actual role sessions after configuration changes, not merely dropdown labels.

4.5 Methods, sources, and limits

Evidence was retained by run and archive batch: TASKs, REPORTs, sessions, tool receipts, chat and public progress, approvals, issues, logs, and EVAL. Qwen and Cursor have late supplements. ZIP integrity, file lists, and hashes were checked before comparison. Complete private logs are not published here.

Timing ends at successful formal final-report submission; termination is separately labelled. The 100-point rubric is a transparent judgment-based assessment with six weights, not a statistical confidence interval. One formal sample per scheme, fixed order, different provider paths, and Cursor's intervention limit generalization. Initialization was shared, but every cache, skill, historical reference, and external condition was not proven identical.

A reference answer for the inspection should cover all five areas with scope, reproducible evidence, conclusions, unknowns, worker receipts, and a PM summary. It must be tied to the actual version and configuration, not a permanent answer independent of system state.

The public package contains the article, figures, selected evidence, scoring/timing data, and source explanations. The concept cover is AI-generated, not a scene photograph. Screenshots are original evidence; diagrams explain mechanisms and are not substitutes for execution. All analytical charts and mechanism diagrams have English editions with selectable vector text. Original screenshots retain the Chinese interface as captured, with English reading guides; literal evidence excerpts preserve their source wording. Supporting source documents are labelled by language in the public source guide.

Public source guide · Evidence inventory (Chinese) · Actual excerpts · EVAL source and checks · Kimi error evidence · EVAL code excerpts · Scores · Timing · Stage timeline

This public article and its supporting files are published separately from the private CodeFlowMu source repository. No access to the private repository is required to read the article or the disclosed evidence excerpts.

Last updated: