You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
While investigating the NFS orphan-file issue (#3532) and auditing the port-scanner code, I traced the full NFS stack in e2b and noticed the sandbox NFS layer is pinned to NFSv3. I'd like to understand whether upgrading to NFSv4 is on the roadmap, and share the trade-offs I found in case it's useful context.
Current architecture
There are two distinct NFS hops:
Sandbox VM (envd)
│ NFSv3 — hard-coded in nfsOptions (init.go:359)
▼
Orchestrator ← go-nfs proxy (e2b fork of willscott/go-nfs)
│ NFSv3 or v4.1 — determined by GCP Filestore tier
▼
GCP Filestore
The client-facing hop (orchestrator → sandbox VM) is fixed at NFSv3 via:
// packages/envd/internal/api/init.go:359"nfsvers=3", // nfs proxy is nfs version 3
The orchestrator uses github.com/e2b-dev/go-nfs (a fork of willscott/go-nfs) as the NFS server. willscott/go-nfs is described as "NFSv3 protocol implementation in pure Golang" — it does not implement NFSv4.
Why NFSv3 makes sense today
I found three reasons the current choice is architecturally sound:
1. pause/resume semantics
The mount options include noac,lookupcache=none with the comment:
// disable caching so that pause/resume works correctly
NFSv3 is stateless by design — each RPC is self-contained. After a VM snapshot/resume, the client retries failed operations transparently. NFSv4 maintains session state (clientid, open stateids, delegations, leases). A resumed VM would face expired leases and would need to go through the NFSv4 grace-period / RECLAIM_COMPLETE recovery protocol, which adds significant complexity to the resume path.
2. No suitable Go NFSv4 server library
There is no production-ready NFSv4 server implementation in Go. Building one from scratch (OPEN/CLOSE state machine, byte-range locking, delegation/recall, RPCSEC_GSS) is a substantial undertaking.
3. NFSv4's main wins don't apply here
NFSv4's Compound RPC and client-side caching (delegation) are its primary performance advantages. Both are negated by the noac,lookupcache=none configuration, which forces every operation to the server anyway.
Where NFSv4 could still help
Despite the above, NFSv4 has properties that could be relevant:
NFS silly-rename is NFSv3-specific. The orphan .nfs* file problem (persistent volume: .nfs* orphan files leak storage when sandbox VMs are forcibly killed #3532) exists because NFSv3's stateless model requires client-side rename-before-delete when a file has open fds. NFSv4's stateful OPEN/CLOSE model handles this at the protocol level — the server knows which files are open and can defer the final delete without renaming.
Single port. NFSv4 runs entirely over TCP port 2049, eliminating the separate portmapper (port 111) and mountd dependencies. The current code already runs a custom portmap server (packages/orchestrator/pkg/portmap/); NFSv4 would remove that layer.
Built-in security model. NFSv4 has a richer ACL and security model, which might matter for multi-tenant volume isolation (currently done at the go-nfs chroot layer).
Questions for the team
Is NFSv4 support on the roadmap? Either for the orchestrator proxy layer, the Filestore-to-orchestrator mount, or both?
Would replacing go-nfs with an NFSv4 implementation be considered, or is the plan to stay with NFSv3 on the proxy layer long-term?
Is the noac,lookupcache=none requirement fundamental (i.e. required by pause/resume correctness), or is there a path to enabling caching that would make NFSv4's performance benefits relevant?
Background
While investigating the NFS orphan-file issue (#3532) and auditing the port-scanner code, I traced the full NFS stack in e2b and noticed the sandbox NFS layer is pinned to NFSv3. I'd like to understand whether upgrading to NFSv4 is on the roadmap, and share the trade-offs I found in case it's useful context.
Current architecture
There are two distinct NFS hops:
The client-facing hop (orchestrator → sandbox VM) is fixed at NFSv3 via:
The orchestrator uses
github.com/e2b-dev/go-nfs(a fork ofwillscott/go-nfs) as the NFS server.willscott/go-nfsis described as "NFSv3 protocol implementation in pure Golang" — it does not implement NFSv4.Why NFSv3 makes sense today
I found three reasons the current choice is architecturally sound:
1. pause/resume semantics
The mount options include
noac,lookupcache=nonewith the comment:// disable caching so that pause/resume works correctlyNFSv3 is stateless by design — each RPC is self-contained. After a VM snapshot/resume, the client retries failed operations transparently. NFSv4 maintains session state (clientid, open stateids, delegations, leases). A resumed VM would face expired leases and would need to go through the NFSv4 grace-period /
RECLAIM_COMPLETErecovery protocol, which adds significant complexity to the resume path.2. No suitable Go NFSv4 server library
There is no production-ready NFSv4 server implementation in Go. Building one from scratch (OPEN/CLOSE state machine, byte-range locking, delegation/recall, RPCSEC_GSS) is a substantial undertaking.
3. NFSv4's main wins don't apply here
NFSv4's Compound RPC and client-side caching (delegation) are its primary performance advantages. Both are negated by the
noac,lookupcache=noneconfiguration, which forces every operation to the server anyway.Where NFSv4 could still help
Despite the above, NFSv4 has properties that could be relevant:
.nfs*file problem (persistent volume: .nfs* orphan files leak storage when sandbox VMs are forcibly killed #3532) exists because NFSv3's stateless model requires client-side rename-before-delete when a file has open fds. NFSv4's stateful OPEN/CLOSE model handles this at the protocol level — the server knows which files are open and can defer the final delete without renaming.packages/orchestrator/pkg/portmap/); NFSv4 would remove that layer.go-nfschroot layer).Questions for the team
Is NFSv4 support on the roadmap? Either for the orchestrator proxy layer, the Filestore-to-orchestrator mount, or both?
Would replacing
go-nfswith an NFSv4 implementation be considered, or is the plan to stay with NFSv3 on the proxy layer long-term?Is the
noac,lookupcache=nonerequirement fundamental (i.e. required by pause/resume correctness), or is there a path to enabling caching that would make NFSv4's performance benefits relevant?How does the team think about the silly-rename / orphan-file problem (persistent volume: .nfs* orphan files leak storage when sandbox VMs are forcibly killed #3532) in light of the NFSv3 constraint? Is the cleanup-at-teardown fix (orchestrator/nfsproxy: clean up .nfs* orphan files on sandbox teardown #3533) the intended long-term approach, or is removing the constraint (NFSv4) also considered?
Happy to contribute research or prototyping work if any of this is useful.
/cc @jakubno @dobrac @ValentaTomas @arkamar @tvi @tomassrnka Looking forward to your feedback.