Skip to content
Merged
Show file tree
Hide file tree
Changes from 3 commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions ci/platform-matrix.json
Original file line number Diff line number Diff line change
Expand Up @@ -64,7 +64,7 @@
"status": "deferred",
"prd_priority": "P1",
"ci_tested": false,
"notes": "The PRD marks this platform as P1. Workstation form-factor with NVIDIA GPUs and the same Docker + NVIDIA Container Toolkit + CDI requirements as DGX Spark. The installer detects DGX Station and offers express install with the pinned `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4` recipe, including an approximately 352 GB model download, without follow-up provider, model, policy, or sandbox-name choices. Pass `--station-deepseek` to use `deepseek-ai/DeepSeek-V4-Flash` for a Station demo while retaining the one-confirmation express flow. The flag requires an interactive terminal, and `/dev/tty` must be available when the installer runs through `curl | bash`. For headless setup, select `NEMOCLAW_PROVIDER=install-vllm` and `NEMOCLAW_VLLM_MODEL=deepseek-v4-flash` instead. Direct managed-vLLM onboarding still defaults to `deepseek-ai/DeepSeek-V4-Flash` when no model override is set. The full NemoClaw onboarding path, including the express recipe, has not been validated end-to-end on physical DGX Station hardware and remains `deferred` until that run is signed off."
"notes": "The PRD marks this platform as P1. Workstation form-factor with NVIDIA GPUs and the same Docker + NVIDIA Container Toolkit + CDI requirements as DGX Spark. Station remains Deferred. On a clean generic Ubuntu 24.04 ARM64 control, the installer detects DGX Station and offers express install with the pinned `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4` recipe, including an approximately 352 GB model download, without follow-up provider, model, policy, or sandbox-name choices. No DGX OS or NVIDIA BaseOS release is currently qualified, so those stock images stop before host preparation. DGX OS 7.5 passed host and plain-container GPU checks, managed inference, tool use, and restart/reuse, but CUDA initialization segfaulted inside the real OpenShell sandbox. Track [NVIDIA/OpenShell#2343](https://github.com/NVIDIA/OpenShell/issues/2343) for requalification. Pass `--station-deepseek` to use `deepseek-ai/DeepSeek-V4-Flash` on the generic Ubuntu control while retaining the one-confirmation express flow. The flag requires an interactive terminal, and `/dev/tty` must be available when the installer runs through `curl | bash`. For headless setup on the generic Ubuntu control, select `NEMOCLAW_PROVIDER=install-vllm` and `NEMOCLAW_VLLM_MODEL=deepseek-v4-flash` instead. Direct managed-vLLM onboarding still defaults to `deepseek-ai/DeepSeek-V4-Flash` when no model override is set."
},
{
"name": "NVIDIA RTX (consumer and Pro workstation GPUs)",
Expand Down Expand Up @@ -147,7 +147,7 @@
"name": "Local vLLM (managed install/start)",
"status": "caveated",
"endpoint_type": "Local OpenAI-compatible",
"notes": "Appears by default on DGX Spark and DGX Station. DGX Station remains deferred until its managed-vLLM onboarding path is validated end-to-end on physical hardware. Generic Linux NVIDIA GPU hosts require `NEMOCLAW_EXPERIMENTAL=1` or `NEMOCLAW_PROVIDER=install-vllm`. Host must have the NVIDIA Container Toolkit installed and a CDI spec present (`onboard` asserts CDI presence). NemoClaw pins runtime images to immutable digests. DGX Spark and DGX Station models without a model-specific runtime use the `linux/arm64` digest published under `nvcr.io/nvidia/vllm:26.05.post1-py3`; generic Linux NVIDIA GPU hosts use the matching `linux/arm64` or `linux/amd64` digest published under `nvcr.io/nvidia/vllm:26.03.post1-py3`. The DGX Station express installer selects `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4` with a pinned Hugging Face revision and the multi-platform index digest published under `vllm/vllm-openai:v0.22.0`; that index resolves to the `linux/arm64` manifest on Station. Direct managed-vLLM profile defaults are listed in `src/lib/inference/vllm-models.ts`: DGX Spark uses `nvidia/Qwen3.6-35B-A3B-NVFP4`, DGX Station uses `deepseek-ai/DeepSeek-V4-Flash`, and Linux NVIDIA GPU uses `nvidia/NVIDIA-Nemotron-3-Nano-4B-FP8`. Image pulls from `nvcr.io` require NGC registry login (`docker login nvcr.io`); onboard prompts for the NGC API key when authentication is missing."
"notes": "Appears by default on DGX Spark and on DGX Station running the generic Ubuntu 24.04 ARM64 control. DGX Station remains Deferred, and no stock DGX OS release is currently qualified. DGX OS 7.5 completed managed-vLLM inference, tool use, and restart/reuse, but CUDA initialization segfaulted inside the real OpenShell sandbox; follow [NVIDIA/OpenShell#2343](https://github.com/NVIDIA/OpenShell/issues/2343). Generic Linux NVIDIA GPU hosts require `NEMOCLAW_EXPERIMENTAL=1` or `NEMOCLAW_PROVIDER=install-vllm`. Host must have the NVIDIA Container Toolkit installed and a CDI spec present (`onboard` asserts CDI presence). NemoClaw pins runtime images to immutable digests. DGX Spark and DGX Station models without a model-specific runtime use the `linux/arm64` digest published under `nvcr.io/nvidia/vllm:26.05.post1-py3`; generic Linux NVIDIA GPU hosts use the matching `linux/arm64` or `linux/amd64` digest published under `nvcr.io/nvidia/vllm:26.03.post1-py3`. The DGX Station express installer selects `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4` with a pinned Hugging Face revision and the multi-platform index digest published under `vllm/vllm-openai:v0.22.0`; that index resolves to the `linux/arm64` manifest on Station. Direct managed-vLLM profile defaults are listed in `src/lib/inference/vllm-models.ts`: DGX Spark uses `nvidia/Qwen3.6-35B-A3B-NVFP4`, DGX Station uses `deepseek-ai/DeepSeek-V4-Flash`, and Linux NVIDIA GPU uses `nvidia/NVIDIA-Nemotron-3-Nano-4B-FP8`. Image pulls from `nvcr.io` require NGC registry login (`docker login nvcr.io`); onboard prompts for the NGC API key when authentication is missing."
}
],

Expand Down
14 changes: 11 additions & 3 deletions docs/get-started/prerequisites.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -41,9 +41,16 @@ If the group change is not active in the current shell, the installer exits with
If you choose the native Linux Ollama install path, the onboard wizard also requires `zstd` for Ollama archive extraction.
The installer also requires `strings` from `binutils` to verify the OpenShell binary before it continues with OpenShell install work.

### DGX Station GB300 OS boundary

On a DGX Station GB300 running the generic Ubuntu 24.04 ARM64 image, accepting express install prepares the host with NVIDIA open driver `610.43.02`, Docker CE `29.6.1` with Buildx, and NVIDIA Container Toolkit `1.19.1`.
DGX OS, NVIDIA BaseOS images, and other Station generations are outside this automatic preparation boundary and stop before host preparation.
On those systems, set `NEMOCLAW_PROVIDER` or `NEMOCLAW_NO_EXPRESS=1` explicitly to continue without Station host automation.
No DGX OS or NVIDIA BaseOS release is currently qualified for this automatic preparation boundary; those images and other Station generations stop before host preparation.
DGX OS 7.5 passed host and plain-container GPU checks, managed Nemotron Ultra inference, tool use, and restart/reuse, but CUDA initialization segfaulted inside the real OpenShell sandbox.
The equivalent sandbox proof passed on the generic Ubuntu control; follow [NVIDIA/OpenShell#2343](https://github.com/NVIDIA/OpenShell/issues/2343) for the blocker and requalification status.
There is no validated in-place conversion from DGX OS.
Back up the Station and use your organization's approved provisioning workflow for a clean generic Ubuntu 24.04 ARM64 image; do not substitute NVIDIA's DGX OS-on-Ubuntu customization procedure, which produces a different kernel and software stack from the validated control.
On those systems, explicitly select a hosted or remote provider, or set `NEMOCLAW_NO_EXPRESS=1` and choose one interactively, to continue without Station host automation.
These overrides do not qualify `install-vllm` or another managed local Station inference path on DGX OS.
The preparation probes package and runtime state first, reuses exact matches, and installs only missing pinned packages, including the NVIDIA Container Toolkit libraries and `nvidia-ctk` CLI.
It permits only the reviewed factory transition from `dkms` `3.0.11-1ubuntu13` to `1:3.4.0-1ubuntu1`.
After reboot, preparation enables NVIDIA's packaged CDI refresh path and service, requires the `nvidia.com/gpu=all` device, and verifies it with a real container launch.
Expand All @@ -61,7 +68,8 @@ After changing pinned packages, the installer exits with status `10`; reboot, si

<Warning title="DGX Station Support Status">
DGX Station remains Deferred.
Full NemoClaw onboarding with this recipe has not completed end-to-end validation on physical DGX Station hardware.
Stock DGX OS is a known no-go for the express path until the OpenShell sandbox CUDA blocker is fixed and the full qualification is repeated.
The generic Ubuntu 24.04 ARM64 configuration remains the validated control.
</Warning>

<Warning title="Docker Group Access">
Expand Down
22 changes: 14 additions & 8 deletions docs/get-started/quickstart.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -121,13 +121,16 @@ Use these details when your first-run path needs more control.

DGX Spark, Station GB300 hosts running the generic Ubuntu 24.04 ARM64 image, and Windows WSL offer an interactive express-install path that chooses a managed local inference option for the platform.
On that Station configuration, accepting the express prompt selects the pinned `nemotron-3-ultra-550b-a55b` managed-vLLM recipe and completes onboarding without more provider, model, policy, or sandbox-name choices.
DGX OS, NVIDIA BaseOS images, and other Station generations stop before host preparation.
Set `NEMOCLAW_PROVIDER` or `NEMOCLAW_NO_EXPRESS=1` explicitly to continue on an unqualified Station without host automation.
No DGX OS or NVIDIA BaseOS release is currently qualified; those images and other Station generations stop before host preparation.
DGX OS 7.5 failed CUDA initialization inside the real OpenShell sandbox even though host, plain-container, managed inference, tool-use, and restart checks passed; follow [NVIDIA/OpenShell#2343](https://github.com/NVIDIA/OpenShell/issues/2343).
There is no validated in-place conversion; use your organization's approved clean generic Ubuntu 24.04 ARM64 provisioning workflow.
Explicitly select a hosted or remote provider, or set `NEMOCLAW_NO_EXPRESS=1` and choose one interactively, to continue without Station host automation.
These overrides do not qualify `install-vllm` or another managed local Station inference path on DGX OS.
It probes the pinned driver, Docker, Buildx, and NVIDIA Container Toolkit versions, reuses exact matches, installs missing pins, and permits only the reviewed factory `dkms` transition.
It establishes NVIDIA CDI through the packaged refresh service, then proves both CDI and `--gpus all` with real container launches.
If the packaged refresh fails or does not advertise `nvidia.com/gpu=all`, preparation prints service diagnostics and stops for administrator repair instead of generating CDI configuration directly.
After it changes pinned packages, the installer exits with status `10` at the required reboot boundary; reboot, sign in, and run the printed exact-commit command to resume the accepted recipe without another prompt.
This automation does not change Station's Deferred support status; physical end-to-end validation remains open.
Station remains Deferred, and stock DGX OS remains outside express install until the sandbox CUDA blocker is fixed and the full qualification is repeated.
Pass `--station-deepseek` to use DeepSeek V4 Flash for a Station demo instead.
The flag selects the interactive express prompt and requires terminal access.
Refer to [Platform Support](../reference/platform-support) and [Choose an Inference Provider](../inference/learn-and-choose/choose-inference-provider) for the current platform behavior.
Expand Down Expand Up @@ -199,18 +202,21 @@ Use these details when your first-run path needs more control.
DGX Spark uses managed vLLM with `qwen3.6-35b-a3b-nvfp4` by default.
DGX Station express install explicitly selects `nemotron-3-ultra-550b-a55b` instead of the Station managed-vLLM profile default, `deepseek-v4-flash`, and discloses the approximately `352 GB` model download before confirmation.
Before onboarding, the Station path requires Station GB300 with the generic Ubuntu 24.04 ARM64 image and checks for NVIDIA open driver `610.43.02`, Docker CE `29.6.1` with Buildx, and NVIDIA Container Toolkit `1.19.1`.
DGX OS, NVIDIA BaseOS images, and other Station generations are outside this automatic preparation boundary.
Set `NEMOCLAW_PROVIDER` or `NEMOCLAW_NO_EXPRESS=1` to continue without Station host automation.
No DGX OS or NVIDIA BaseOS release is currently qualified; those images and other Station generations are outside this automatic preparation boundary.
DGX OS 7.5 failed CUDA initialization inside the real OpenShell sandbox even though host, plain-container, managed inference, tool-use, and restart checks passed; follow [NVIDIA/OpenShell#2343](https://github.com/NVIDIA/OpenShell/issues/2343).
There is no validated in-place conversion; use your organization's approved clean generic Ubuntu 24.04 ARM64 provisioning workflow.
Explicitly select a hosted or remote provider, or set `NEMOCLAW_NO_EXPRESS=1` and choose one interactively, to continue without Station host automation.
These overrides do not qualify `install-vllm` or another managed local Station inference path on DGX OS.
Preparation reuses exact versions, installs missing pinned packages, permits only the reviewed `dkms` transition from `3.0.11-1ubuntu13` to `1:3.4.0-1ubuntu1`, and refuses every other mismatched version.
It requires the packaged NVIDIA CDI refresh service to advertise `nvidia.com/gpu=all`, and verifies CDI and `--gpus all` with real container launches.
If the packaged refresh fails or omits that device, preparation prints service diagnostics and stops for administrator repair instead of generating CDI configuration directly.
If NVIDIA Docker runtime registration or a post-change acceptance probe fails, preparation restores the prior Docker daemon configuration.
Changing pinned packages exits with status `10` for a reboot; after you sign in and run the printed exact-commit command, the accepted express recipe resumes without another prompt.
This automation does not change Station's Deferred support status; physical end-to-end validation remains open.
Station remains Deferred, and stock DGX OS remains outside express install until the sandbox CUDA blocker is fixed and the full qualification is repeated.
To select DeepSeek V4 Flash while retaining the one-confirmation Station express flow, run `curl -fsSL https://www.nvidia.com/nemoclaw.sh | bash -s -- --station-deepseek`.
The `--station-deepseek` flag requires an interactive terminal; in a `curl | bash` pipeline, `/dev/tty` must be available.
Without terminal access, the installer stops before it installs Docker or build dependencies instead of ignoring the flag.
For a headless or CI install on a prepared DGX Station, omit the flag and select the same managed-vLLM recipe explicitly.
For a headless or CI install on a prepared DGX Station running the qualified generic Ubuntu 24.04 ARM64 control, omit the flag and select the same managed-vLLM recipe explicitly.
Comment thread
coderabbitai[bot] marked this conversation as resolved.
Outdated

```bash
curl -fsSL https://www.nvidia.com/nemoclaw.sh | \
Expand All @@ -229,7 +235,7 @@ Use these details when your first-run path needs more control.

<Warning>
Express install automates the configuration but does not change DGX Station's Deferred support status.
Full NemoClaw onboarding with this recipe has not completed end-to-end validation on physical DGX Station hardware.
Stock DGX OS 7.5 failed qualification at CUDA initialization inside the OpenShell sandbox; the generic Ubuntu 24.04 ARM64 configuration remains the control.
</Warning>

The installer auto-launches `nemoclaw onboard` when it can find the new binary.
Expand Down
Loading
Loading