Skip to main content

([arXiv][1])---# 但还有第二种 memory 浪费Guest 里面可能:

  1. ([arXiv][1])

    ---

    # 但还有第二种 memory 浪费

    Guest 里面可能:

    
    page
    page
    page
    page
    


    已经不用了。

    但是 guest 不告诉 host:

    > “这些 page 我不要了。”

    于是 host 认为:

    
    这 VM 还在用。
    


    ---

    # DSec:DAMON + balloon free-page reporting

    大概:

    
    DAMON
     ↓
    identify cold pages
     ↓
    reclaim
     ↓
    balloon reports free pages
     ↓
    host gets memory back
    


    实验:

    
    time-integrated memory
    ↓
    21.2%
    


    而:

    
    DAX
    +
    DAMON/FPR
    


    组合效果最好。([arXiv][1])

    ---

    # 第九课:CPU 又怎么办?

    现在:

    
    1000 sandbox
    


    很多都 idle。

    所以:

    
    CPU overcommit
    


    很合理。

    但是问题来了。

    假设:

    
    Sandbox A = latency sensitive
    Sandbox B = best effort
    


    如果两个线程落在:

    
    same physical core
    SMT sibling
    


    即使:

    
    A priority high
    B priority low
    


    B 还是可能影响 A。

    ---

    # DSec 的两层 QoS

    ### 第一层

    
    BE → SCHED_IDLE
    


    即:

    > 有 LS 工作的时候,BE 让路。

    ### 第二层

    
    Linux Core Scheduling
    


    防止:

    
    LS thread
    +
    BE thread
    


    跑在同一个 physical core 的 sibling threads 上。

    实验:

    
    无 QoS:
    latency +45.2%
    
    SCHED_IDLE + Core Scheduling:
    latency +17.3%
    


    ([arXiv][1])

    ---

    # 第十课:这才是我认为 DSec 最重要的设计

    现在假设:

    
    RL training job
    


    正在运行。

    GPU:

    
    ████████████████████
    


    Sandbox:

    
    ████████████████████
    


    突然 GPU cluster 需要抢占这个 training job。

    传统设计:

    
    GPU pod
     ├── model server
     ├── RL framework
     └── agent loop
    
    GPU pod killed
           ↓
    agent loop killed
    


    但是:

    
    Sandbox
    


    可能还活着。

    于是:

    
    Sandbox state ≠ Agent loop state
    


    灾难。

    ---

    # 以前 DeepSeek 的办法

    保存:

    
    command log
    


    例如:

    
    1. cat README
    2. grep foo
    3. edit foo.py
    4. pytest
    5. ...
    


    恢复:

    
    sandbox restored
           +
    replay command log
    


    但是有一个大问题:

    > command 不一定 idempotent。

    例如:

    
    mkdir foo
    


    第一次成功。

    Replay:

    
    mkdir foo
    


    可能失败。

    更糟糕:

    
    curl -X POST ...
    


    Replay 会产生**重复 side effect**。

    所以他们不得不:

    
    replay
    +
    recorded result
    +
    avoid re-execution
    


    系统复杂度很高。([arXiv][1])

    ---

    # V4.1 之后的关键变化

    DeepSeek 把:

    
    Agent Loop
    


    从 GPU training pod 里面拿出来。

    变成:

    
                 DSec
                  │
           ┌──────┴──────┐
           │             │
    Agent Sandbox    Worker Container
           │             │
           └──────┬──────┘
                  │
            complete rollout state
    


    然后:

    
    GPU Training
         │
         │ preempt
         ▼
       gone
    
    DSEC
         │
         ├── agent state
         ├── sandbox state
         └── rollout state
              ↓
           preserved
    


    GPU job 恢复以后:

    
    GPU training
         │
         ▼
     reconnect
         │
         ▼
    continue rollout
    


    不需要重新 replay。

    论文把这描述成:

    > worker container + agent sandbox jointly retain complete rollout state and act as the single source of truth.

    ([arXiv][1])

    ---

    # 这意味着什么?

    这是 DSec 真正从:

    > Sandbox Service

    升级成:

    > Agent Runtime Infrastructure

    的地方。

    因为它不只是:

    
    run command
    


    而是:

    
    preserve agent execution state
    


    ---

    # 第十一课:Sandbox Pause/Resume

    但是又来了一个问题。

    GPU 被抢占:

    
    100,000 sandboxes
    


    不能全部继续吃 RAM。

    所以:

    
    sandbox state
        ≠
    sandbox active memory
    


    这两个东西必须分开。

    ---

    # Container

    Pause:

    
    docker pause
           ↓
    freeze process tree
           ↓
    memory.swap.max
           ↓
    memory.reclaim
           ↓
    reclaim memory
    


    Resume:

    
    MADV_WILLNEED
           ↓
    prefetch
           ↓
    docker unpause
    


    ([arXiv][1])

    ---

    # MicroVM

    更直接:

    
    running VM
        ↓
    snapshot memory + execution state
        ↓
    kill Firecracker process
        ↓
    RAM released
    


    恢复:

    
    snapshot
       ↓
    new Firecracker process
       ↓
    restore
       ↓
    continue
    


    ([arXiv][1])

    ---

    # 所以这里出现了一个非常漂亮的抽象

    你可以把 DSec Sandbox 理解成:

    
                    Sandbox
                       │
              ┌────────┴────────┐
              │                 │
           Logical state      Runtime
              │                 │
              │              CPU/RAM
              │
              ▼
          persistent
              │
              │
              └───────────────┐
                              │
                     pause / resume
                              │
                              ▼
                      different runtime
    


    也就是说:

    > Sandbox 是一个 stateful logical machine,而不是永远绑定一个进程。

    这和传统 container service 的思想已经非常不一样。