Meet OceanBase AI Lakebase, the unified database for operational data, real-time analytics, and AI.

Explore ->

Meet OceanBase AI Lakebase, the unified database for operational data, real-time analytics, and AI. Explore ->

PowerContext: A Foundation for Human–Agent Collaboration

Qing Tang
Qing Tang
Published on October 10, 2026Updated on 2026-10-10
20 minute read
Key Takeaways
  • PowerContext is an open-source (Apache 2.0) context runtime for human–agent collaboration. It turns goals, decisions, verified progress, blockers, and next steps into Memory and Handoff records that persist across sessions, models, and agent hosts such as Codex and Claude Code.
  • It targets the gap that checkpoints and memory retrieval leave open: when a new session, agent, or person takes over, the reasons behind rejected approaches and the list of still-unverified assumptions usually stay behind in the old conversation.
  • Handoffs carry provenance, revisions, and explicit acceptance, and context is assembled within a byte budget, so the next participant can check what they inherit instead of trusting a summary. You can install it with uv and connect Codex in a few commands.

The AI-Native Organization’s Dilemma: If Every Agent Is Capable, Why Does Work Still Break Down?

At 9:40 p.m. on a Friday, a message lands in the on-call group: payment callbacks are intermittently creating duplicate ledger entries in the canary environment. In plain terms, the system is recording the same payment twice. You grab your laptop from the couch, open the company project, and start a new Codex session.

The conversation from earlier that day is already closed. It had grown too long, mixing changes in product requirements, two failed retry strategies, a rejected distributed lock approach, and three sets of logs from a tester. The new session is clean and can read the repository again. But it does not know three things scattered elsewhere: the lock approach discussed during the day would slow reconciliation beyond what the customer could accept; reproducing the issue requires replaying requests with the same idempotency key; and the customer’s environment has no cache cluster, so the duplicate entries must be fixed using the existing database.

You spend 20 minutes explaining the background again. The agent then recommends the same lock approach from earlier that day. You have to reject it again. To finish before the canary window closes, you bring in two more agents: one to assess the risks of the database changes, and another to check compatibility with older clients. All three can read the repository and return results quickly, but each has only part of the background. One includes a compensation job that is out of scope for tonight, another reviews a different code revision, and the third still assumes the customer has a cache cluster. You wanted to split the work across agents. Instead, you have another round of coordination and rework.

The code preserves the implementation. The logs preserve what happened during execution. But the judgments that explain how the team arrived here remain in the old conversation. A new agent can see the files, yet a person still has to explain the connections between them: why one approach was rejected, why a particular test is the one that counts, and why this customer is an exception.

Scale this up to an AI-native organization, and the gaps multiply as more roles get involved. Consider an enterprise login feature. A product manager first defines the scope in one agent. Engineering switches to another and discovers that the customer’s environment cannot reach the external login service. Sales uses a third agent to build a customer demo. A developer advocate finally opens another agent to write the quickstart. Every role moves faster, but the work as a whole may not. Product scope stays in a conversation, failure details are buried in terminal output, customer constraints sit in the CRM, and the person writing the documentation has the latest code but does not know which capabilities can already be promised to customers.

As execution capacity grows, teams increasingly need a way for the next participant to build on the thinking and validation already completed. The foundation discussed in this article is a layer of work context independent of any single conversation, model, or tool. Organized around the work in progress, it preserves the current goal, the reasoning behind decisions, and the state of progress, so the work can continue as people and agents change places.

The Emerging Consensus in 2026: The Agent Bottleneck Is Shifting from Execution to Collaboration

A range of public examples in 2026 are approaching this problem from different directions: tasks are getting longer, more participants are involved, and the infrastructure needed to maintain continuity is becoming more important.

  • January: Organizational context becomes part of an agent’s capabilities. OpenAI’s internal data agent uses table schemas, human annotations, code, organizational knowledge, memory, and runtime state together. Knowing how to write SQL is only part of the job; the model also needs to know how the company defines a metric and which data can be combined in a calculation. OpenAI: Inside OpenAI’s in-house data agent
  • February: State starts to extend beyond a single request. OpenAI and AWS announce a stateful runtime environment for production tasks spanning steps, tools, approvals, and system state. Tracking how far work has progressed and recovering after failures become responsibilities of the runtime environment. OpenAI: Stateful Runtime Environment for Agents
  • March: Progress records across sessions become essential support for long tasks. Anthropic describes an approach to sustained scientific work and revisits Claude’s compiler-building effort across roughly 2,000 sessions. The scientific computing example maintains a dedicated progress file covering the current state, validation results, failed approaches, and the reasons they failed, so later sessions do not revisit dead ends. Anthropic: Long-running Claude for scientific computing
  • April: Human context switching becomes a bottleneck. OpenAI’s Symphony team observes that engineers can usually manage three to five sessions comfortably at once. The team responds by organizing work around tasks and deliverables, allowing people to focus on reviewing results. This is an observation from the team’s practice, not a fixed limit for everyone. OpenAI: Symphony
  • May: Context starts to be managed independently. LangChain introduces Context Hub to store, version, collaborate on, and publish context files. Around the same time, Google Cloud describes checkpoints, recovery, and human approval mechanisms for long tasks. These address content management and execution continuity respectively, showing that writing a good prompt is only one part of the work. LangChain: Introducing LangSmith Context Hub; Google Cloud: Production-ready AI agents
  • June: Long tasks and parallel work begin to show their scale. In OpenAI’s research, 70.2% of sampled individual users had submitted tasks estimated to require more than an hour of human work. Internally, users at the 99th percentile of usage often generated more than 60 hours of parallel agent runtime per day. The first figure relies on model estimates of human effort; the second measures parallel runtime. Neither can be treated directly as hours saved. Microsoft also introduces context capabilities at Build that connect people, email, documents, meetings, and business data. OpenAI: How agents are transforming work; Microsoft Build 2026
  • August: Collaboration reveals failure modes of its own. Anthropic’s multi-agent experiments observe code merge conflicts, isolated ways of working, and collective errors caused by agents converging on the same decisions. Adding participants does not automatically produce reliable teamwork. Anthropic: Patterns and problems in emerging multiagent systems

Taken together, these examples point to one conclusion: as agents take on more execution, coordination, verification, and decision-making account for a larger share of the overall work. They also remind us that running long tasks, managing context, and enabling teamwork are distinct problems. A larger chat window cannot solve them all.

Where Collaboration Starts: Let Work Leave the Conversation

Complex tasks usually cross several boundaries: a conversation window fills up, a model changes, an engineer finishes their shift, or a task moves from development to testing, presales, or the on-call team. At every boundary, the team has to answer the same questions again.

  • What needs to be accomplished now?
  • Which results have been verified, and where is the evidence?
  • Which approaches were tried and abandoned, and why?
  • Which questions remain unanswered, and what should the next participant do first?

Conversation history, Memory, and Handoff all relate to context, but they solve different problems.

  • Execution continuity answers, “How does the same session keep running?” A resume feature, checkpoint, or transcript lets the same host recover its previous state. This is useful, but the state is usually tied to the original tool and session format. Switch to another host, and that execution state does not travel with you.
  • Knowledge continuity answers, “What should still be remembered in the future?” Memory and RAG are well suited to retrieving facts, constraints, and decisions with lasting value, such as “The customer’s environment has no cache cluster” or “Compensation jobs are out of scope for this release.” This information may be useful again, but it does not necessarily explain how far the current work has progressed.
  • Work continuity answers, “How should the next participant continue from here?” Handoff organizes the current goal, verified progress, blockers, next steps, and evidence for whoever takes over. That may be a new session, a new model, another agent host, or a person.

A usable foundation for collaboration must distinguish these three forms of continuity. Execution state belongs with the host. Judgments that remain valuable across tasks belong in Memory. Work in progress passes to the next participant through Handoff. Saving everything together only makes the conversation history longer; it does not make taking over easier.

The distinction can be captured in two sentences:

Memory answers, “What should still be remembered in the future?”

Handoff answers, “How should the next participant continue from here?”

MCP connects agents to tools and data. A2A lets agents discover and invoke other agents. Checkpoints let a run recover. These address connection, invocation, and recovery, but they do not define the state in which unfinished work should be passed to the next participant when it leaves its original session. Teams need a layer of work context that belongs to no single session or host.

Where We Diverge: Three Choices Behind PowerContext

The industry has no shortage of memory systems. What is missing is a way for work to continue. That is why we evolved PowerMem, our earlier open-source memory layer for agents, into PowerContext. PowerContext is a context runtime layer for collaboration between people and agents. It does not duplicate knowledge bases, issue trackers, code repositories, or observability systems. It organizes the evidence, judgments, state, and next steps worth carrying forward from those systems into project context that can persist across sessions, models, and agents.

Memory keeps facts from being forgotten. Handoff keeps work from starting over.

Three choices shape this approach.

First, the product is organized around work, not isolated pieces of information. PowerContext focuses on goals, key judgments, verified progress, blockers, next steps, and the relationships between them.

Second, context must serve the handoff. More context is not automatically better, and context is not another name for a complete log. Its value depends on whether the next participant can understand faster, make fewer guesses, and continue more reliably. Good context should let someone see where to pick up within a few minutes, without rereading the entire conversation.

Third, the next participant may be an agent or a person. PowerContext does not seek to remove people from the loop. As agents become more autonomous, people need an even clearer view of key judgments and the ability to take over and change course. Anthropic’s work on trustworthy agents likewise identifies maintaining human control, preserving transparency, and knowing when to return decisions to people as core principles. (Trustworthy agents in practice) If it is to serve that work, PowerContext must answer three questions in sequence: Can someone pick it up and continue? Can they trust what is handed over? Can they take it beyond a single agent?

PowerContext’s Answer: Work You Can Pick Up, Trust, and Take with You

These three questions form an ordered product flow, rather than a feature checklist. First, the work must become a Handoff that the next participant can act on. That Handoff must have traceable sources and clear boundaries. Finally, it must be able to follow the project across sessions, participants, and agents.

Ready to Continue: Let Handoff Carry Unfinished Work

Traditional memory asks, “What information from the past is worth recalling in the future?” PowerContext takes the next step: How can unfinished work be continued? A Handoff should tell the next participant how to act and where to pause for confirmation. The incident at the beginning of this article could be organized like this:

Goal                  Fix duplicate payment callback entries before the canary window closes
Scope                 Fix duplicate entries tonight; assess compensation jobs separately
Verified              Replaying requests with the same idempotency key reliably reproduces the issue
Abandoned             The distributed lock approach from earlier today: reconciliation latency
                      exceeds what the customer can accept
Customer constraints  No cache cluster on site; the solution must work with the existing database
To verify             Impact of database changes on production; compatibility with older clients
Next step             Complete the database risk review and compatibility checks, then let a
                      person decide whether to proceed with the canary rollout
Evidence              Reproduction logs, environment notes, the corresponding code revision,
                      and test conditions

“To verify” needs to be as visible as “Verified.” If a handoff preserves only good news, the next participant may treat unverified parts as settled conclusions too. The purpose of this work package is to show where to continue and which evidence is still needed before doing so.

The delegated scope first makes the goal and completion criteria explicit. A person or agent then organizes the goal, progress, blockers, next steps, and supporting evidence into a work package ready for inspection. When a milestone needs to be preserved, the package can be committed as a traceable revision. The receiver checks it against the current repository, current instructions, and their own capabilities and permissions, then accepts, requests clarification, or declines. When the task finishes, becomes blocked, or is interrupted, the outcome and verification evidence are recorded.

Two distinctions are easy to miss. Sending a handoff does not mean anyone has accepted it. Accepting a handoff does not mean the task is complete. Recording preparation, acceptance, and outcome separately lets the team see exactly where the work stands. Acceptance itself grants no additional permissions; publishing still requires the applicable authorization.

A handoff serves the current task by default. It does not automatically become long-term memory. The reason the lock approach was rejected should travel with this incident response. A customer environment constraint deserves a separate record only if it remains relevant to later tasks. Keeping temporary arrangements and lasting knowledge in their proper places helps later agents avoid reading “Do this tonight” as “Always do this.”

People and agents can also read different views of the same revision. People see business impact, tradeoffs, and pending decisions first. Agents can expand execution details, evidence, and verification steps. The level of detail may differ, but the references should point to the same verifiable materials.

Being able to pick up the work is only the first step. An incorrect, outdated, or unverifiable handoff may be more dangerous than no handoff at all.

Trustworthy: Give Context Traceable Sources and Clear Boundaries

“Retrieve more history for the model” sounds reasonable, but it can also amplify outdated judgments, incorrect summaries, malicious content, and irrelevant noise. In the on-call scenario, speculation in the group chat, customer environment notes, and test logs carry different evidentiary weight. A test may be valid only for a specific code revision and set of conditions. A constraint for one customer cannot automatically become a product-wide rule. Compressing all of this into “Resolved” loses the boundaries the next participant most needs to understand. PowerContext therefore brings provenance, lifecycle management, and runtime budgets into the same flow:

External work materials
    ↓
Source: preserve the original provenance and the limits of the evidence
    ↓
Memory / Artifact: create context assets that can be revised or deactivated,
                   with their history preserved
    ↓
PreparedContext: assemble at request time based on scope, relevance, and byte budget
    ↓
Agent: use retrieved content as untrusted history with citations

Four product principles matter most here:

  • Capturing does not mean trusting. Capturing a prompt creates a Source. It does not turn a casual user remark directly into long-term Memory.
  • Updates do not erase history. Revising or deactivating a Memory preserves historical revisions. As with Git commits, changes to each piece of information remain traceable.
  • Relevance does not justify unlimited injection. In PowerContext, the server assembles context within relevance constraints and a byte budget. By default, each request is limited to 8,000 UTF-8 bytes. If retrieval fails, the system degrades gracefully without blocking the user’s original work.
  • Self-improvement stays under control. An agent or caller can submit an Experience or Skill candidate. Only after it passes Review does it become an immutable Experience or Skill.

These boundaries are especially important in enterprise projects. The system needs to remember more than customer preferences: scope changes, interface constraints, accepted risks, past commitments, and unresolved issues all matter. When this information is scattered across CRM records, tickets, email, and meeting notes, retrieving a relevant snippet is not enough to determine whether it belongs to the current project or has been superseded by a newer decision. Scope, revision, and citation make “Which project?”, “Which decision?”, and “Where is the evidence?” first-class information. Likewise, a successful attempt should not automatically become a rule for the entire team. A command that works in one repository may cause damage in another environment. PowerContext separates observing a practice from approving it as a team asset: first create a candidate supported by evidence, then have a person review it. Only more stable, actionable procedures become Experience and, in turn, Skills that can be explicitly exported. At this point, context can be inspected and governed. But if it remains in an agent’s private history, even trustworthy context is still an island.

Portable: Change Agents without Rebuilding Project Context

Real teams do not use the same model, framework, IDE, or agent forever.

You may debug in Codex today and review in Claude Code tomorrow. One team builds business agents with LangChain; another orchestrates workflows with LangGraph. Local development needs to stay lightweight, while team deployments need shared access controls, persistence, and observability. PowerContext therefore provides the Server as an independent runtime layer:

  • Local development can use OceanBase seekdb directly as the underlying data store; team deployments can use an OceanBase cluster.
  • A common set of capabilities is exposed through HTTP/OpenAPI and Streamable HTTP MCP.
  • The Runtime maintains project scopes, revisions that preserve history, citations, and the Handoff and Review contracts.
  • Agent integrations capture, retrieve, and present context at the right moments, without maintaining separate context stores of their own.

The project currently provides 13 integrations: Codex, Claude Code, DeepSeek Harness, ZCode, Hermes Agent, Pi Coding Agent, OpenClaw, OpenCode, WorkBuddy, Bub, Pydantic AI, LangChain, and LangGraph. They all serve the same purpose: agents can change while the team retains control of its project context. Work is no longer tied to a single session or a vendor’s private history. With the handoff object, trust boundaries, and a shared layer across agents established, we can follow a piece of work through its lifecycle to see how they connect.

How a Piece of Work Moves through PowerContext

PowerContext’s main flow can be condensed into five steps. It builds connections between evidence and work state on top of existing systems, rather than moving all information into a new one.

  • Step 1: Bring in the evidence. Code, documents, tickets, traces, agent trajectories, and human input can all become Sources. PowerContext focuses on source references and relationships between evidence; it does not replace the user’s existing data systems.
  • Step 2: Preserve judgments worth reusing. Decisions, constraints, results, and state can be recorded explicitly. With a generation model configured, candidate information can also be extracted from Sources. Every Memory retains its provenance, and later revisions or deactivation do not erase its history.
  • Step 3: Assemble bounded context when it is needed. Before the agent receives a request, the Runtime generates PreparedContext based on project scope, relevance, and budget, with citations attached. The goal is to provide the minimum sufficient context for the next step, rather than everything that might be relevant. This aligns with Anthropic's direction on just-in-time context and progressive disclosure: keep lightweight references and expand them as needed at runtime, instead of filling the window in advance.
  • Step 4: Create a Handoff before participants change. The goal, verified progress, blockers, next steps, and evidence are organized into a work package for a new session, task, model, agent host, or person.
  • Step 5: Use feedback from reuse without bypassing governance. Practices validated across multiple tasks enter the Experience / Skill Candidate and Review process. Approved assets then support future work, creating a feedback loop for reuse under human control.

Conceptually, this completes the loop. Whether it actually reduces repeated work and improves task outcomes still needs to be answered through evaluation.

Does It Actually Improve Results? Let Public Benchmarks Answer

Context products can easily fall into two forms of self-validation: showing a demo that looks clever, or using the number of retrieved items as a substitute for actual task value. PowerContext measures two things: the accuracy and cost of long-term memory question answering, and the task completion rate of a real coding agent.

LoCoMo: Finding the Right Context Fast, Without Stuffing in the Full History

LoCoMo is a public benchmark for long-conversation memory. The public evaluation protocol for PowerContext’s end-to-end pipeline covers 10 conversations, 272 sessions, 5,882 dialogue turns, and 1,986 questions. Of these, 1,540 questions in categories 1–4 are included in scoring. The results are shown below:

The chart presents three sets of key results:

SystemQA accuracySearch p95 latencyAnswer tokens / question
PowerContext90.78%1.38 s~1.65K
PowerMem87.79%1.44 s~0.9K
Full-context baseline52.9%17.12 s~26K

Compared with the full-context baseline, PowerContext improves accuracy by 37.88 percentage points, reduces search p95 latency by approximately 91.9%, and uses approximately 93.7% fewer answer tokens per question. Compared with PowerMem, accuracy rises by 2.99 percentage points and search p95 latency is slightly lower. PowerContext also uses more answer tokens than PowerMem. We include that tradeoff in the chart because a good context system seeks an explainable balance between accuracy, latency, and cost, rather than the lowest possible value for any one metric. The 90.78% result corresponds to 1,398 correct answers out of 1,540 questions and makes no claim about LoCoMo’s event summarization or multimodal dialogue generation tasks. See the benchmark methodology and limitations for details.

SWE-bench Pro Public v2: Context Must Ultimately Improve Task Outcomes

Memory question answering measures whether information can be recalled. Coding agent evaluation comes closer to measuring whether the work can actually be completed. PowerContext ran an OFF / ON comparison across 731 SWE-bench Pro public v2 tasks in Codex. Both groups used gpt-5.6-sol with reasoning effort set to medium:

  • PowerContext OFF: 602 / 731, a task resolution rate of 82.35%.
  • PowerContext ON: 634 / 731, a task resolution rate of 86.73%.
  • Difference: +4.38 percentage points and +32 tasks.

As an end-to-end result, this addresses the most important question: PowerContext’s value can be tested against completion rates on real agent tasks, beyond simply counting how many memories it has saved. Two caveats apply. This is a paired run on a pinned task set, not an official SWE-bench Pro leaderboard submission, and because agent runs are stochastic, the scores describe these two runs only. The comparison also measures the effect of the whole system and does not isolate the contribution of Handoff.

Try It Yourself: Get Started in 3 Minutes

Local installation requires Git, Python 3.11 or later, and the uv package manager. The following example uses PowerContext 1.2.0:

uv tool install "powercontext[cli,server]==1.2.0"
powercontext setup codex --ref powercontext-v1.2.0
powercontext server run

Use matching versions of the PowerContext tool and agent integrations. To try capabilities from the current development branch, install with:

uv tool install --force "powercontext[cli,server] @ git+https://github.com/oceanbase/powercontext.git@master"
powercontext setup codex --source oceanbase/powercontext --ref master
powercontext setup claude-code --source oceanbase/powercontext --ref master
powercontext server run

Once the server is running, check it from another terminal:

powercontext doctor
powercontext doctor codex

The local setup stores context in a database by default and provides a dashboard. For handoffs across tools, connect both sides to the same service and explicitly select the same project Scope. Connecting to the same service does not mean they have selected the same project context. See the Agent Integration Guide for the complete steps. Then try three small things:

  1. Explicitly ask the agent to save a judgment with lasting value.
  2. Open a new session and check whether it can retrieve and use that judgment correctly.
  3. While a task is still unfinished, say Handoff this work, then have the receiver continue from that specific handoff revision.

Note: Manually recording memories and preparing and committing handoffs does not require a generation model to be configured on the PowerContext server. Automatic memory extraction, more comprehensive retrieval, and self-improvement capabilities require the corresponding configuration.

Why Open Source: Context Infrastructure Must Be Inspectable, Extensible, and Portable

Context influences an agent’s judgments and carries a team’s project history. Infrastructure like this is difficult to trust over the long term if it is available only as a hosted black box, with no way to inspect data contracts or verify retrieval boundaries. PowerContext is open source under the Apache License 2.0. The repository includes the SDK, Server, canonical OpenAPI contract, agent integrations, RFCs, tests, and the LoCoMo and SWE-bench Pro evaluation toolchains.

The project is developing a shared scope for community contributions:

  • Contribute new agent integrations to make project context available in more places where work happens.
  • Improve Memory extraction, hybrid search, reranking, PreparedContext, and Handoff.
  • Contribute reproducible evaluations, beyond demos that merely appear to work.
  • Participate in RFC design and propose improvements to the governance semantics of scope, evidence, Review, Experience, and Skill.
  • Improve documentation and contribute real deployment feedback and failure cases, so the product’s documented boundaries stay more credible than its marketing.

This is what community contribution means for PowerContext: jointly defining how people and agents share, hand off, and continue work, rather than filling plugin gaps in a closed product.

Conclusion and Outlook: What Becomes Scarce Is Not Better Answers, but Work That Doesn’t Get Lost

Conclusion

Models will keep improving, context windows will keep growing, and more agents will run at the same time for longer periods. We will also have more memory systems, workflows, tools, traces, and protocols. But an organization’s lasting, compounding value comes from preserving verified judgments so that those who follow do not have to start from zero, rather than from the volume of content it has generated.

When one agent hands a task to another, when an agent returns an uncertain decision to a person, when an engineer switches hosts, or when an incident passes to the next shift, what needs to travel is work context with evidence, clear boundaries, and a next step, rather than the entire history. This is the context infrastructure PowerContext aims to build. The goal goes beyond helping agents remember more: it is to ensure that the work people and agents have done together is not easily lost and can be continued whenever needed.

Roadmap

PowerContext’s roadmap centers on one goal: let context accumulate as work progresses, make it easier for people and agents to continue each other’s work, and keep improving through real outcomes. Building on the existing Memory, Handoff, Experience, and Skill capabilities, we will focus on six areas:

  1. Observability: Make context capture, retrieval, and use visible, and assess their impact on task outcomes, time, and cost.
  2. Self-improvement: Distill experience from task results and user feedback, then use validated lessons to improve Experience and Skill continuously.
  3. The handoff experience: Simplify handoff and resumption so work can continue more smoothly across sessions, agents, and people.
  4. Context quality: Improve information selection, conflict resolution, and management of outdated content so accumulated context stays accurate, relevant, and usable.
  5. Ecosystem and integrations: Deepen integrations with major agents and development frameworks, reducing the effort required for installation, configuration, and everyday use.
  6. Team collaboration and production readiness: Strengthen knowledge sharing, access control, auditing, and operations to support reliable, long-term use by teams.

PowerContext is developed in the open on GitHub. The quickest way to see whether it fits your workflow is to connect two agents to the same project Scope, hand off one unfinished task, and check what the receiver still has to ask.

Share
X
linkedin
mail