uv and connect Codex in a few commands.
At 9:40 p.m. on a Friday, a message lands in the on-call group: payment callbacks are intermittently creating duplicate ledger entries in the canary environment. In plain terms, the system is recording the same payment twice. You grab your laptop from the couch, open the company project, and start a new Codex session.
The conversation from earlier that day is already closed. It had grown too long, mixing changes in product requirements, two failed retry strategies, a rejected distributed lock approach, and three sets of logs from a tester. The new session is clean and can read the repository again. But it does not know three things scattered elsewhere: the lock approach discussed during the day would slow reconciliation beyond what the customer could accept; reproducing the issue requires replaying requests with the same idempotency key; and the customer’s environment has no cache cluster, so the duplicate entries must be fixed using the existing database.
You spend 20 minutes explaining the background again. The agent then recommends the same lock approach from earlier that day. You have to reject it again. To finish before the canary window closes, you bring in two more agents: one to assess the risks of the database changes, and another to check compatibility with older clients. All three can read the repository and return results quickly, but each has only part of the background. One includes a compensation job that is out of scope for tonight, another reviews a different code revision, and the third still assumes the customer has a cache cluster. You wanted to split the work across agents. Instead, you have another round of coordination and rework.

The code preserves the implementation. The logs preserve what happened during execution. But the judgments that explain how the team arrived here remain in the old conversation. A new agent can see the files, yet a person still has to explain the connections between them: why one approach was rejected, why a particular test is the one that counts, and why this customer is an exception.
Scale this up to an AI-native organization, and the gaps multiply as more roles get involved. Consider an enterprise login feature. A product manager first defines the scope in one agent. Engineering switches to another and discovers that the customer’s environment cannot reach the external login service. Sales uses a third agent to build a customer demo. A developer advocate finally opens another agent to write the quickstart. Every role moves faster, but the work as a whole may not. Product scope stays in a conversation, failure details are buried in terminal output, customer constraints sit in the CRM, and the person writing the documentation has the latest code but does not know which capabilities can already be promised to customers.
As execution capacity grows, teams increasingly need a way for the next participant to build on the thinking and validation already completed. The foundation discussed in this article is a layer of work context independent of any single conversation, model, or tool. Organized around the work in progress, it preserves the current goal, the reasoning behind decisions, and the state of progress, so the work can continue as people and agents change places.
A range of public examples in 2026 are approaching this problem from different directions: tasks are getting longer, more participants are involved, and the infrastructure needed to maintain continuity is becoming more important.
Taken together, these examples point to one conclusion: as agents take on more execution, coordination, verification, and decision-making account for a larger share of the overall work. They also remind us that running long tasks, managing context, and enabling teamwork are distinct problems. A larger chat window cannot solve them all.
Complex tasks usually cross several boundaries: a conversation window fills up, a model changes, an engineer finishes their shift, or a task moves from development to testing, presales, or the on-call team. At every boundary, the team has to answer the same questions again.

Conversation history, Memory, and Handoff all relate to context, but they solve different problems.
resume feature, checkpoint, or transcript lets the same host recover its previous state. This is useful, but the state is usually tied to the original tool and session format. Switch to another host, and that execution state does not travel with you.A usable foundation for collaboration must distinguish these three forms of continuity. Execution state belongs with the host. Judgments that remain valuable across tasks belong in Memory. Work in progress passes to the next participant through Handoff. Saving everything together only makes the conversation history longer; it does not make taking over easier.
The distinction can be captured in two sentences:
Memory answers, “What should still be remembered in the future?”
Handoff answers, “How should the next participant continue from here?”
MCP connects agents to tools and data. A2A lets agents discover and invoke other agents. Checkpoints let a run recover. These address connection, invocation, and recovery, but they do not define the state in which unfinished work should be passed to the next participant when it leaves its original session. Teams need a layer of work context that belongs to no single session or host.
The industry has no shortage of memory systems. What is missing is a way for work to continue. That is why we evolved PowerMem, our earlier open-source memory layer for agents, into PowerContext. PowerContext is a context runtime layer for collaboration between people and agents. It does not duplicate knowledge bases, issue trackers, code repositories, or observability systems. It organizes the evidence, judgments, state, and next steps worth carrying forward from those systems into project context that can persist across sessions, models, and agents.
Memory keeps facts from being forgotten. Handoff keeps work from starting over.
Three choices shape this approach.
First, the product is organized around work, not isolated pieces of information. PowerContext focuses on goals, key judgments, verified progress, blockers, next steps, and the relationships between them.
Second, context must serve the handoff. More context is not automatically better, and context is not another name for a complete log. Its value depends on whether the next participant can understand faster, make fewer guesses, and continue more reliably. Good context should let someone see where to pick up within a few minutes, without rereading the entire conversation.
Third, the next participant may be an agent or a person. PowerContext does not seek to remove people from the loop. As agents become more autonomous, people need an even clearer view of key judgments and the ability to take over and change course. Anthropic’s work on trustworthy agents likewise identifies maintaining human control, preserving transparency, and knowing when to return decisions to people as core principles. (Trustworthy agents in practice) If it is to serve that work, PowerContext must answer three questions in sequence: Can someone pick it up and continue? Can they trust what is handed over? Can they take it beyond a single agent?
These three questions form an ordered product flow, rather than a feature checklist. First, the work must become a Handoff that the next participant can act on. That Handoff must have traceable sources and clear boundaries. Finally, it must be able to follow the project across sessions, participants, and agents.
Traditional memory asks, “What information from the past is worth recalling in the future?” PowerContext takes the next step: How can unfinished work be continued? A Handoff should tell the next participant how to act and where to pause for confirmation. The incident at the beginning of this article could be organized like this:
Goal Fix duplicate payment callback entries before the canary window closes
Scope Fix duplicate entries tonight; assess compensation jobs separately
Verified Replaying requests with the same idempotency key reliably reproduces the issue
Abandoned The distributed lock approach from earlier today: reconciliation latency
exceeds what the customer can accept
Customer constraints No cache cluster on site; the solution must work with the existing database
To verify Impact of database changes on production; compatibility with older clients
Next step Complete the database risk review and compatibility checks, then let a
person decide whether to proceed with the canary rollout
Evidence Reproduction logs, environment notes, the corresponding code revision,
and test conditions“To verify” needs to be as visible as “Verified.” If a handoff preserves only good news, the next participant may treat unverified parts as settled conclusions too. The purpose of this work package is to show where to continue and which evidence is still needed before doing so.

The delegated scope first makes the goal and completion criteria explicit. A person or agent then organizes the goal, progress, blockers, next steps, and supporting evidence into a work package ready for inspection. When a milestone needs to be preserved, the package can be committed as a traceable revision. The receiver checks it against the current repository, current instructions, and their own capabilities and permissions, then accepts, requests clarification, or declines. When the task finishes, becomes blocked, or is interrupted, the outcome and verification evidence are recorded.
Two distinctions are easy to miss. Sending a handoff does not mean anyone has accepted it. Accepting a handoff does not mean the task is complete. Recording preparation, acceptance, and outcome separately lets the team see exactly where the work stands. Acceptance itself grants no additional permissions; publishing still requires the applicable authorization.
A handoff serves the current task by default. It does not automatically become long-term memory. The reason the lock approach was rejected should travel with this incident response. A customer environment constraint deserves a separate record only if it remains relevant to later tasks. Keeping temporary arrangements and lasting knowledge in their proper places helps later agents avoid reading “Do this tonight” as “Always do this.”
People and agents can also read different views of the same revision. People see business impact, tradeoffs, and pending decisions first. Agents can expand execution details, evidence, and verification steps. The level of detail may differ, but the references should point to the same verifiable materials.
Being able to pick up the work is only the first step. An incorrect, outdated, or unverifiable handoff may be more dangerous than no handoff at all.
“Retrieve more history for the model” sounds reasonable, but it can also amplify outdated judgments, incorrect summaries, malicious content, and irrelevant noise. In the on-call scenario, speculation in the group chat, customer environment notes, and test logs carry different evidentiary weight. A test may be valid only for a specific code revision and set of conditions. A constraint for one customer cannot automatically become a product-wide rule. Compressing all of this into “Resolved” loses the boundaries the next participant most needs to understand. PowerContext therefore brings provenance, lifecycle management, and runtime budgets into the same flow:
External work materials
↓
Source: preserve the original provenance and the limits of the evidence
↓
Memory / Artifact: create context assets that can be revised or deactivated,
with their history preserved
↓
PreparedContext: assemble at request time based on scope, relevance, and byte budget
↓
Agent: use retrieved content as untrusted history with citationsFour product principles matter most here:
These boundaries are especially important in enterprise projects. The system needs to remember more than customer preferences: scope changes, interface constraints, accepted risks, past commitments, and unresolved issues all matter. When this information is scattered across CRM records, tickets, email, and meeting notes, retrieving a relevant snippet is not enough to determine whether it belongs to the current project or has been superseded by a newer decision. Scope, revision, and citation make “Which project?”, “Which decision?”, and “Where is the evidence?” first-class information. Likewise, a successful attempt should not automatically become a rule for the entire team. A command that works in one repository may cause damage in another environment. PowerContext separates observing a practice from approving it as a team asset: first create a candidate supported by evidence, then have a person review it. Only more stable, actionable procedures become Experience and, in turn, Skills that can be explicitly exported. At this point, context can be inspected and governed. But if it remains in an agent’s private history, even trustworthy context is still an island.
Real teams do not use the same model, framework, IDE, or agent forever.

You may debug in Codex today and review in Claude Code tomorrow. One team builds business agents with LangChain; another orchestrates workflows with LangGraph. Local development needs to stay lightweight, while team deployments need shared access controls, persistence, and observability. PowerContext therefore provides the Server as an independent runtime layer:
The project currently provides 13 integrations: Codex, Claude Code, DeepSeek Harness, ZCode, Hermes Agent, Pi Coding Agent, OpenClaw, OpenCode, WorkBuddy, Bub, Pydantic AI, LangChain, and LangGraph. They all serve the same purpose: agents can change while the team retains control of its project context. Work is no longer tied to a single session or a vendor’s private history. With the handoff object, trust boundaries, and a shared layer across agents established, we can follow a piece of work through its lifecycle to see how they connect.
PowerContext’s main flow can be condensed into five steps. It builds connections between evidence and work state on top of existing systems, rather than moving all information into a new one.

PreparedContext based on project scope, relevance, and budget, with citations attached. The goal is to provide the minimum sufficient context for the next step, rather than everything that might be relevant. This aligns with Anthropic's direction on just-in-time context and progressive disclosure: keep lightweight references and expand them as needed at runtime, instead of filling the window in advance.Conceptually, this completes the loop. Whether it actually reduces repeated work and improves task outcomes still needs to be answered through evaluation.
Context products can easily fall into two forms of self-validation: showing a demo that looks clever, or using the number of retrieved items as a substitute for actual task value. PowerContext measures two things: the accuracy and cost of long-term memory question answering, and the task completion rate of a real coding agent.
LoCoMo is a public benchmark for long-conversation memory. The public evaluation protocol for PowerContext’s end-to-end pipeline covers 10 conversations, 272 sessions, 5,882 dialogue turns, and 1,986 questions. Of these, 1,540 questions in categories 1–4 are included in scoring. The results are shown below:

The chart presents three sets of key results:
| System | QA accuracy | Search p95 latency | Answer tokens / question |
| PowerContext | 90.78% | 1.38 s | ~1.65K |
| PowerMem | 87.79% | 1.44 s | ~0.9K |
| Full-context baseline | 52.9% | 17.12 s | ~26K |
Compared with the full-context baseline, PowerContext improves accuracy by 37.88 percentage points, reduces search p95 latency by approximately 91.9%, and uses approximately 93.7% fewer answer tokens per question. Compared with PowerMem, accuracy rises by 2.99 percentage points and search p95 latency is slightly lower. PowerContext also uses more answer tokens than PowerMem. We include that tradeoff in the chart because a good context system seeks an explainable balance between accuracy, latency, and cost, rather than the lowest possible value for any one metric. The 90.78% result corresponds to 1,398 correct answers out of 1,540 questions and makes no claim about LoCoMo’s event summarization or multimodal dialogue generation tasks. See the benchmark methodology and limitations for details.
Memory question answering measures whether information can be recalled. Coding agent evaluation comes closer to measuring whether the work can actually be completed. PowerContext ran an OFF / ON comparison across 731 SWE-bench Pro public v2 tasks in Codex. Both groups used gpt-5.6-sol with reasoning effort set to medium:

As an end-to-end result, this addresses the most important question: PowerContext’s value can be tested against completion rates on real agent tasks, beyond simply counting how many memories it has saved. Two caveats apply. This is a paired run on a pinned task set, not an official SWE-bench Pro leaderboard submission, and because agent runs are stochastic, the scores describe these two runs only. The comparison also measures the effect of the whole system and does not isolate the contribution of Handoff.
Local installation requires Git, Python 3.11 or later, and the uv package manager. The following example uses PowerContext 1.2.0:
uv tool install "powercontext[cli,server]==1.2.0"
powercontext setup codex --ref powercontext-v1.2.0
powercontext server runUse matching versions of the PowerContext tool and agent integrations. To try capabilities from the current development branch, install with:
uv tool install --force "powercontext[cli,server] @ git+https://github.com/oceanbase/powercontext.git@master"
powercontext setup codex --source oceanbase/powercontext --ref master
powercontext setup claude-code --source oceanbase/powercontext --ref master
powercontext server runOnce the server is running, check it from another terminal:
powercontext doctor
powercontext doctor codexThe local setup stores context in a database by default and provides a dashboard. For handoffs across tools, connect both sides to the same service and explicitly select the same project Scope. Connecting to the same service does not mean they have selected the same project context. See the Agent Integration Guide for the complete steps. Then try three small things:
Handoff this work, then have the receiver continue from that specific handoff revision.Note: Manually recording memories and preparing and committing handoffs does not require a generation model to be configured on the PowerContext server. Automatic memory extraction, more comprehensive retrieval, and self-improvement capabilities require the corresponding configuration.
Context influences an agent’s judgments and carries a team’s project history. Infrastructure like this is difficult to trust over the long term if it is available only as a hosted black box, with no way to inspect data contracts or verify retrieval boundaries. PowerContext is open source under the Apache License 2.0. The repository includes the SDK, Server, canonical OpenAPI contract, agent integrations, RFCs, tests, and the LoCoMo and SWE-bench Pro evaluation toolchains.
The project is developing a shared scope for community contributions:
This is what community contribution means for PowerContext: jointly defining how people and agents share, hand off, and continue work, rather than filling plugin gaps in a closed product.
Models will keep improving, context windows will keep growing, and more agents will run at the same time for longer periods. We will also have more memory systems, workflows, tools, traces, and protocols. But an organization’s lasting, compounding value comes from preserving verified judgments so that those who follow do not have to start from zero, rather than from the volume of content it has generated.
When one agent hands a task to another, when an agent returns an uncertain decision to a person, when an engineer switches hosts, or when an incident passes to the next shift, what needs to travel is work context with evidence, clear boundaries, and a next step, rather than the entire history. This is the context infrastructure PowerContext aims to build. The goal goes beyond helping agents remember more: it is to ensure that the work people and agents have done together is not easily lost and can be continued whenever needed.
PowerContext’s roadmap centers on one goal: let context accumulate as work progresses, make it easier for people and agents to continue each other’s work, and keep improving through real outcomes. Building on the existing Memory, Handoff, Experience, and Skill capabilities, we will focus on six areas:
PowerContext is developed in the open on GitHub. The quickest way to see whether it fits your workflow is to connect two agents to the same project Scope, hand off one unfinished task, and check what the receiver still has to ask.

AI era doesn't need another heavy, complex enterprise database. It needs agility. It needs flexibility. We went back to the drawing board to understand what an AI application actually needs from a database. Our answer is OceanBase seekdb


On the DABstep Global Leaderboard, OceanBase DataPilot agent has secured the top spot, maintaining a significant lead over the runner-up for a month. The secret to our SOTA results was a fundamental shift in engineering paradigm: moving from "Prompt Engineering" to "Asset Engineering."


OceanBase HTAP runs TP and AP workloads in one cluster using row, column, or hybrid storage, cutting analytics latency from T+1 to near real time.
