Runtime, workflow, recovery, tools, skills, observability and governance mechanisms.
SWE-bench and its Verified subset demonstrate that issue clarity, test validity, environment reproducibility, and human adjudication are part of the engineering system being evaluated, not peripheral dataset maintenance.
OpenAI Agents SDK distinguishes manager-style specialist calls from handoffs that transfer active control, showing that multi-agent design must model ownership and authority rather than treating every delegation as the same tool call.
LangGraph, OpenHands, CrewAI, and AutoGen show a shared shift from short-lived agent loops toward persistent state, controlled interruption, recovery, sandboxing, and structured runtime operations.
OpenHands, CrewAI, AutoGen, and LangGraph show that reusable agent capability is moving from hidden prompt text into explicit skills, plugins, tools, workflows, message contracts, and observable events.
The design decision that turns the joinwell52 repository into a bilingual, versioned and continuously updated public research portal.
An engineering benchmark for skills, connection health, automation triggers, runtime options and operator experience.