Replies: 1 comment
|
A concrete use case for this request is evaluating coding models through native T3 delegation while keeping reference solutions and grader data inaccessible to the child. Devin/SWE-2 is one example; the delegation contract makes this relevant across providers. The parent should retain evaluation controls and receive task status/results. The child should see only the supplied fixture and approved tools. A separate Git worktree, a starting directory, or read-only permissions do not establish that boundary. Verified current behaviorI checked public
These checks establish current core behavior, not live answer access by a model. Provider-native filesystem sandbox enforcement was not tested. The shared projection check used Codex, Claude, Cursor, Grok, OpenCode, Antigravity and Devin targets; their native sandboxes should not be assumed equivalent. Devin's host-terminal path is a specific additional boundary to contain. Latest stable Smallest useful directionCould delegated tasks target an explicitly isolated execution environment while retaining For an evaluation profile, useful negative checks would be:
A separate OS-isolated T3 environment with fresh application data is a possible external workaround, but I have not tested it. The request here is to keep native delegation and its lifecycle while making that boundary explicit and verifiable. Related: #15184 asks for unattended read-only children; evaluation isolation additionally requires restricted reads. #16188 asks for project-scoped MCP access; evaluation children also need isolation from selected data and parent transcripts within the same project. Reduced offline reproductionSave the following as import assert from 'node:assert/strict';
import vm from 'node:vm';
const base='https://raw.githubusercontent.com/pingdotgg/t3code/758dc290e900e35958d37fec41fed4a357ad60fa/';
const files=['apps/server/src/orchestration-v2/SubagentProjection.ts', 'apps/server/src/provider/acp/AcpClientPolicy.ts', 'apps/server/src/provider/acp/AcpClientTerminals.ts'];
const sources=new Map(await Promise.all(files.map(async file=>{
const response=await fetch(base+file);assert(response.ok,`${file}: ${response.status}`);
return [file,await response.text()] as const;
})));
const transpiler = new Bun.Transpiler({loader:'ts', target:'bun'});
function load(file: string, name: string, context: vm.Context) {
const source = transpiler.transformSync(sources.get(file)!);
const start = source.indexOf(`function ${name}(`);
assert(start >= 0, name);
const end = source.indexOf('\n}', start);
assert(end > start);
const fragment = source.slice(start, end + 2);
vm.runInContext(fragment, context, {timeout:1000});
}
const context = vm.createContext({ChildProcess:{make:(command: string,args: string[],options: unknown)=>({command,args,options})}});
load('apps/server/src/orchestration-v2/SubagentProjection.ts','makeSubagentChildThread',context);
const parent={id:'parent',projectId:'project',worktreePath:'/fixture-parent',branch:'main',historyOrigin:'parent-history',lineage:{rootThreadId:'parent'}};
const providers=['codex','claudeAgent','cursor','grok','opencode','antigravity','acpRegistry_devin'];
for (const providerInstanceId of providers) {
const child=context.makeSubagentChildThread({parentThread:parent,childThreadId:'child',parentNodeId:'node',activeProviderThreadId:null,providerInstanceId,modelSelection:{model:'synthetic'},title:'fixture',now:'synthetic',createdBy:'agent',creationSource:'mcp'});
assert.equal(child.worktreePath,parent.worktreePath);assert.equal(child.projectId,parent.projectId);assert.equal(child.branch,parent.branch);assert.equal(child.historyOrigin,undefined);
}
const policy='apps/server/src/provider/acp/AcpClientPolicy.ts';
for (const name of ['unknownRecord','isAcpReadKind','acpReadDisposition','acpPolicyRequiresApproval','isAcpMutationKind','acpOperationDisposition','acpClientExecuteDisposition']) load(policy,name,context);
const policies=[{runtimeMode:'full-access'},{runtimeMode:'approval-required'},{runtimeMode:'auto-accept-edits'},{runtimeMode:'approval-required',approvalPolicy:'never',sandboxPolicy:{type:'readOnly'}}];
const rows=policies.map(p=>{
const read=context.acpOperationDisposition({...p,cwd:'/fixture'},{kind:'read',locations:[{path:'/outside/sentinel.txt'}]});
assert.equal(read,'allow');
return {policy:p,read,execute:context.acpClientExecuteDisposition(p)};
});
assert.equal(rows[0].execute,'allow');assert.equal(rows[3].execute,'deny');
load('apps/server/src/provider/acp/AcpClientTerminals.ts','acpTerminalCommand',context);
const cmd=context.acpTerminalCommand({request:{command:'/bin/cat',args:['/outside/sentinel.txt'],cwd:'/outside'},defaultCwd:'/fixture',shellCommands:true});
assert.equal(cmd.options.cwd,'/outside');assert.equal(cmd.args[0],'/outside/sentinel.txt');
console.log(JSON.stringify({ref:'758dc290e900e35958d37fec41fed4a357ad60fa',providersWithInheritedWorkspace:providers,policyResults:rows,terminalOutsideCwdPreserved:true,commandsExecuted:0,modelCalls:0},null,2));The result lists the seven inherited-workspace targets, |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Motivation
Orchestrator V2 makes the parent-thread model extremely compelling for long-running project work: a parent can decompose a goal, launch multiple workers, choose providers/models, monitor them, steer them, and aggregate their results.
One of the main things preventing me from using T3 as the primary control plane for this workflow is compute isolation and elasticity.
Today, multiple delegated workers generally share the compute of the parent environment. For larger tasks, I'd like the parent to be able to fan work out across isolated, disposable cloud environments.
For example:
From the user's perspective, the interaction should remain entirely inside the parent T3 thread:
The orchestrator can then decide whether it needs one worker or five without requiring me to manually create and switch between environments.
This would provide something similar to the elastic execution model of cloud coding agents / Cursor Projects, while keeping T3 both agent-provider-agnostic and compute-provider-agnostic.
Possible approaches
I can see two ways this could fit T3's existing architecture.
1. Allow workers to target another T3 Environment
A delegated/thread launch could optionally target an existing environment:
Provisioning could remain completely outside T3.
For example, an MCP/tool could:
T3 itself wouldn't need provider-specific infrastructure code. Its responsibility would remain:
2. Keep the normal T3 worker local and let it supervise cloud compute
Another option may require even less coupling to T3's Environment model.
The Orchestrator could continue creating completely normal V2 child threads:
Each T3 worker would act as a supervisor and use an MCP/tool to:
This has the nice property that T3 retains its existing native Orchestrator UX:
The actual compute remains an implementation detail behind the worker.
This approach may already be possible almost entirely outside T3. If so, perhaps the only useful additions to T3 would be small primitives that make externally supervised workers feel more native.
Compute providers
I don't think T3 should need to own integrations for every compute vendor.
Ideally the architecture remains generic enough that any of these could provide the disposable compute:
The recently proposed Docker Sandboxes integration seems like an excellent first implementation because it already maps naturally onto T3's SSH Environment abstraction.
Longer term, though, keeping orchestration separate from compute provisioning would allow users to choose whichever provider fits their cost, isolation, or infrastructure requirements.
Why this matters
I'm currently using T3 as my main coding-agent interface.
For me, this is one of the main remaining gaps between:
and
Orchestrator V2 already provides most of the coordination model.
Adding a clean path for workers to consume isolated, elastic compute would allow a single parent thread to coordinate many parallel coding agents without requiring one increasingly large permanent machine.
The important part for me isn't a specific sandbox vendor. It's preserving this separation:
That would make T3 a very compelling open and provider-agnostic alternative to vertically integrated cloud-agent environments.
All reactions