Agentic AI OS
│
┌───────────┼───────────┐
│ │ │
Compute Storage Security
│ │ │
FnCall/VM EROFS/3FS eBPF/AppArmor
│ │ │
└───────────┼───────────┘
│
Agent Runtime
│
RL integration
│
State / Pause / Resume
它实际上负责:
compute abstraction
resource scheduling
storage abstraction
state management
security boundary
network policy
lifecycle management
failure recovery
elasticity
所以它已经非常接近:
> Agent Infrastructure Operating System
---
# 但是这里有一个非常重要的“反直觉点”
DSEC 最厉害的地方**不是任何单一技术**。
不是:
EROFS
也不是:
Firecracker
也不是:
3FS
也不是:
Power-of-k
而是:
# 它把 Agentic RL 的 workload 特征直接反推成了系统架构。
对应关系非常漂亮:
| Agent workload 特征 | DSec mechanism |
| ------------------------- | --------------------------------- |
| bursty creation | Power-of-k + elastic placement |
| sparse CPU | CPU overcommit |
| long-lived state | persistent sandbox |
| memory expensive | DAX + reclamation |
| heterogeneous isolation | FnCall / Container / MicroVM / VM |
| huge images | composable layers |
| low image reuse | shared 3FS |
| only partial image access | on-demand EROFS |
| small writes | local writable layer |
| RL preemption | pause/resume |
| rollout state | DSec-owned agent loop |
| reward hacking | AppArmor + eBPF |
| peak overflow | cloud bursting |
这张表其实就是**整篇论文的骨架**。
---
# 最后,用一个完整例子把它串起来
假设 DeepSeek 要训练:
> 32,000 个 Coding Agents
任务:
> 修复 32,000 个不同 repository。
---
## Step 1:RL Framework
32K tasks
│
▼
libdsec.create()
---
## Step 2:Placement
API
↓
IAM
↓
Placement
↓
Power-of-k
↓
Edge
---
## Step 3:Environment
每个 sandbox:
Ubuntu base
+
repo layer
+
DeepSeek Harness layer
+
toolkit layer
+
local writable layer
---
## Step 4:启动
不是:
pull 6GB
unpack 6GB
而:
local metadata
+
EROFS mount
+
3FS on-demand data
---
## Step 5:Agent 开始工作
LLM
↓
shell
↓
sandbox
↓
stdout
↓
LLM
↓
edit
↓
sandbox
↓
pytest
整个过程中:
CPU = bursty
RAM = persistent
---
## Step 6:CPU overcommit
1000 sandbox
↓
most idle
↓
share CPU
---
## Step 7:Memory pressure
cold pages
↓
DAMON
↓
reclaim
MicroVM:
DAX
↓
avoid duplicated page cache
---
## Step 8:GPU training 被抢占
GPU job
↓
PREEMPT
Dsec:
sandbox
↓
pause
↓
memory reclaim
但:
agent state
sandbox state
仍然存在。
---
## Step 9:GPU 恢复
GPU trainer
↓
reconnect
↓
Dsec worker
↓
resume sandbox
↓
continue agent
不用 replay:
cat
grep
edit
pytest
...
---
## Step 10:Reward
pytest passed
↓
reward
↓
RL update
---
# 现在我希望你真正掌握的“DSEC 心智模型”
如果让我把整个论文压缩成 **7句话**:
> 1. Agentic RL 的核心不是“让 LLM 生成 token”,而是让 LLM 在真实环境中持续行动。
> 2. 因此,每条 rollout 都需要一个 stateful、isolated、可执行的 sandbox。
> 3. RL 又会一次性产生几千到几万个 sandbox,所以传统 Docker/VM 的启动、镜像分发和资源管理模式无法直接扩展。
> 4. DSec 通过多种 sandbox backend、弹性 placement、CPU overcommit、memory reclamation 和 stateful lifecycle,把几十万 sandbox 放进一个集群。
> 5. EROFS + 3FS 则解决了最大的 storage bottleneck:不要把整个 image 搬过来,只在真正访问时读取真正需要的数据。
> 6. 最重要的是,Dsec 把 Agent rollout state 从 GPU training job 中解耦出来,使 RL preemption 不再意味着 rollout 丢失。
> 7. 所以 DSec 本质上不是一个 Sandbox Runtime,而是一个为 Agentic RL 设计的 Elastic Agent Execution Platform。
---
## 以及一个非常关键的边界
这里我要帮你把论文里“事实”和“我的理解”分开。
论文的实验并**没有**证明:
> “DSec 让 Agent 能力提升了多少”。
论文的实验主要验证: