Skip to main content

Agentic AI OS

  1. Agentic AI OS
                       │
           ┌───────────┼───────────┐
           │           │           │
       Compute       Storage     Security
           │           │           │
     FnCall/VM     EROFS/3FS    eBPF/AppArmor
           │           │           │
           └───────────┼───────────┘
                       │
                 Agent Runtime
                       │
                RL integration
                       │
              State / Pause / Resume
    


    它实际上负责:

    
    compute abstraction
    resource scheduling
    storage abstraction
    state management
    security boundary
    network policy
    lifecycle management
    failure recovery
    elasticity
    


    所以它已经非常接近:

    > Agent Infrastructure Operating System

    ---

    # 但是这里有一个非常重要的“反直觉点”

    DSEC 最厉害的地方**不是任何单一技术**。

    不是:

    
    EROFS
    


    也不是:

    
    Firecracker
    


    也不是:

    
    3FS
    


    也不是:

    
    Power-of-k
    


    而是:

    # 它把 Agentic RL 的 workload 特征直接反推成了系统架构。

    对应关系非常漂亮:

    | Agent workload 特征 | DSec mechanism |
    | ------------------------- | --------------------------------- |
    | bursty creation | Power-of-k + elastic placement |
    | sparse CPU | CPU overcommit |
    | long-lived state | persistent sandbox |
    | memory expensive | DAX + reclamation |
    | heterogeneous isolation | FnCall / Container / MicroVM / VM |
    | huge images | composable layers |
    | low image reuse | shared 3FS |
    | only partial image access | on-demand EROFS |
    | small writes | local writable layer |
    | RL preemption | pause/resume |
    | rollout state | DSec-owned agent loop |
    | reward hacking | AppArmor + eBPF |
    | peak overflow | cloud bursting |

    这张表其实就是**整篇论文的骨架**。

    ---

    # 最后,用一个完整例子把它串起来

    假设 DeepSeek 要训练:

    > 32,000 个 Coding Agents

    任务:

    > 修复 32,000 个不同 repository。

    ---

    ## Step 1:RL Framework

    
    32K tasks
       │
       ▼
    libdsec.create()
    


    ---

    ## Step 2:Placement

    
    API
     ↓
    IAM
     ↓
    Placement
     ↓
    Power-of-k
     ↓
    Edge
    


    ---

    ## Step 3:Environment

    每个 sandbox:

    
    Ubuntu base
    +
    repo layer
    +
    DeepSeek Harness layer
    +
    toolkit layer
    +
    local writable layer
    


    ---

    ## Step 4:启动

    不是:

    
    pull 6GB
    unpack 6GB
    


    而:

    
    local metadata
          +
    EROFS mount
          +
    3FS on-demand data
    


    ---

    ## Step 5:Agent 开始工作

    
    LLM
     ↓
    shell
     ↓
    sandbox
     ↓
    stdout
     ↓
    LLM
     ↓
    edit
     ↓
    sandbox
     ↓
    pytest
    


    整个过程中:

    
    CPU = bursty
    RAM = persistent
    


    ---

    ## Step 6:CPU overcommit

    
    1000 sandbox
         ↓
    most idle
         ↓
    share CPU
    


    ---

    ## Step 7:Memory pressure

    
    cold pages
     ↓
    DAMON
     ↓
    reclaim
    


    MicroVM:

    
    DAX
     ↓
    avoid duplicated page cache
    


    ---

    ## Step 8:GPU training 被抢占

    
    GPU job
       ↓
    PREEMPT
    


    Dsec:

    
    sandbox
     ↓
    pause
     ↓
    memory reclaim
    


    但:

    
    agent state
    sandbox state
    


    仍然存在。

    ---

    ## Step 9:GPU 恢复

    
    GPU trainer
          ↓
    reconnect
          ↓
    Dsec worker
          ↓
    resume sandbox
          ↓
    continue agent
    


    不用 replay:

    
    cat
    grep
    edit
    pytest
    ...
    


    ---

    ## Step 10:Reward

    
    pytest passed
       ↓
    reward
       ↓
    RL update
    


    ---

    # 现在我希望你真正掌握的“DSEC 心智模型”

    如果让我把整个论文压缩成 **7句话**:

    > 1. Agentic RL 的核心不是“让 LLM 生成 token”,而是让 LLM 在真实环境中持续行动。

    > 2. 因此,每条 rollout 都需要一个 stateful、isolated、可执行的 sandbox。

    > 3. RL 又会一次性产生几千到几万个 sandbox,所以传统 Docker/VM 的启动、镜像分发和资源管理模式无法直接扩展。

    > 4. DSec 通过多种 sandbox backend、弹性 placement、CPU overcommit、memory reclamation 和 stateful lifecycle,把几十万 sandbox 放进一个集群。

    > 5. EROFS + 3FS 则解决了最大的 storage bottleneck:不要把整个 image 搬过来,只在真正访问时读取真正需要的数据。

    > 6. 最重要的是,Dsec 把 Agent rollout state 从 GPU training job 中解耦出来,使 RL preemption 不再意味着 rollout 丢失。

    > 7. 所以 DSec 本质上不是一个 Sandbox Runtime,而是一个为 Agentic RL 设计的 Elastic Agent Execution Platform。

    ---

    ## 以及一个非常关键的边界

    这里我要帮你把论文里“事实”和“我的理解”分开。

    论文的实验并**没有**证明:

    > “DSec 让 Agent 能力提升了多少”。

    论文的实验主要验证: