Skip to main content
AI

NVIDIA OpenShell on AWS: Your Agent, Your Rules

A model connects to an agent inside an OpenShell policy boundary for files, processes, and network access
Conceptual flow · model ↔ agent → toolsMotion off: reduced-motion preference

The model decides what to do. The agent uses tools to carry out those decisions. NVIDIA OpenShell supplies a policy-controlled execution environment. Software Sushi packages that runtime for AWS, with a repeatable image recipe, setup, and customer guidance. Together, these pieces give an agent a workspace with explicit rules for files, processes, and network access.

OpenShell is not an AI model. You choose the agent and model service separately. This article explains the boundaries, the Software Sushi AMI, and a real model-backed task we ran inside a sandbox.

Four pieces, four different jobs

NVIDIA describes OpenShell as a runtime that enforces policies around agent actions. A useful way to understand it is to separate the intelligence, the tool user, and the place where those tools execute. See NVIDIA's OpenShell overview for the upstream product description.

01 · Model

Decides the next action

Generates responses and proposes tool calls. It can run at a hosted provider or on a private model server.

02 · Agent

Uses the tools

Connects the model to a workflow: reading files, running commands, calling services, and returning results for the next decision.

03 · OpenShell

Controls execution

Runs the workload in a sandbox with policy boundaries for file access, processes, and network connections.

Conceptual view · intelligence and execution

Your model service

Hosted API or private model server. Receives requests and returns responses.

Agent inside OpenShell

Its tools work within the configured policy. Access to the model endpoint must also be allowed.

FilesProcessesNetwork
Conceptual diagram: the model can be outside the sandbox while the agent's tools run inside it. Network policy must permit the connection.

Watch the one-minute introduction

The short introduction connects these four roles before the practical walkthroughs. All three videos use English narration and captions. Watch here or on Software Sushi’s YouTube channel; text transcripts accompany each video.

Introduction · 58 seconds · model, agent, OpenShell, and Software Sushi. Watch on YouTube.
Read the introduction transcript

00:00–00:13 · Give your agent a controlled workspace. Your AI model decides what to do. An agent carries out those decisions. OpenShell gives the agent a workspace with rules for files, commands, and network access.

00:13–00:29 · How the pieces fit together. The model provides intelligence. The agent uses tools. NVIDIA OpenShell applies the execution boundaries. Software Sushi packages that runtime for AWS, with setup, guides, and a consistent image recipe.

00:29–00:44 · Watch a real task. In the practical video, a real model chooses tools to read a small CSV, write a report, and verify the result. We also attempt one harmless write outside the allowed workspace and show the system blocking it.

00:44–00:58 · Start your own workflow. Start an instance, connect, create a sandbox, and run your workflow. Choose your model service separately. Use the hands-on video and its source files to see exactly how this example works.

From an AMI to a working sandbox

An Amazon Machine Image is a reusable starting image for an EC2 instance. It contains the software needed to boot the machine; it does not execute an agent by itself.

AMI · The template

A prepared installation

The recipe fixes the runtime versions, installation, configuration, and checks used to produce an image.

EC2 · Your machine

An instance in your account

Launching the image creates a machine with its own resources, storage, permissions, and lifecycle.

Sandbox · The workspace

A workload with rules

OpenShell's gateway manages sandbox creation and access. You supply the workload and the policy for its tools.

One AMI can launch several independent instances. Publishing a newer image does not automatically update machines already running. The tested candidate uses a CPU instance for the agent runtime; the model in our demonstration runs on a separate machine.

Watch the full terminal walkthrough

This recorded 3-minute-26-second walkthrough starts with a fresh instance from the private AMI. The installed OpenShell runtime started and its smoke checks passed without replacing runtime files. You can see the real openshell term dashboard, sandbox setup, five model generations, a verified CSV report, the outside-write denial, and scoped cleanup.

Recorded OpenShell terminal dashboard showing a healthy Sushi gateway before creating a sandbox. Watch the full terminal walkthrough on YouTube · 3:26
Private AMI test · actual ANSI output from a remote terminal session, rendered as a terminal. Idle waits are shortened and final frames are held for narration; the original sessions are preserved. English narration, captions, and nine chapters are available on YouTube.
Read the full terminal walkthrough transcript

00:00–00:20 · What you will do. Let's use OpenShell on AWS, packaged by Software Sushi. In this walkthrough, we connect to an instance, create a controlled workspace, and let an AI agent turn a small CSV into a report. Then we check a permission boundary and clean up. You will see the actual terminal and its live results.

00:20–00:42 · Start and connect. Start with an instance from the OpenShell image and your approved AWS access. The connection in this example uses Systems Manager, so there is no public SSH port to open. We first check that the service started automatically and that the gateway responds. Keep credentials and account details out of your recording.

00:42–01:03 · Watch the live dashboard. This is OpenShell's terminal dashboard. It gives you a live view of the environment, including the gateway and your sandboxes. The panel is useful for watching a workspace become ready and inspecting its state. OpenShell supplies the controlled runtime. The model and the agent that uses it are separate pieces.

01:03–01:27 · Create a bounded workspace. Here we create one sandbox for the example. It has one CPU and two hundred and fifty-six megabytes of memory. Its policy allows work in the sandbox directory and one reviewed model route. We wait for the configuration to settle before starting the agent. These limits belong to this demonstration, and you can choose appropriate limits for your own workload.

01:27–01:51 · Give the agent a task. The input is a CSV with three tasks. Our small Python agent asks a real model to choose from four reviewed tools. Watch the output as the model reads the CSV, writes a Markdown report, and checks it against the original data. The model serves the requests from an existing private connection; it is not installed inside this AMI.

01:51–02:15 · Inspect the real result. Now open the report the agent actually wrote. It lists the three tasks and adds their durations to sixty seconds. That number comes from the sample data; it is not a performance benchmark. The independent check compares each row and the total against the CSV. This gives us an observable result, rather than just a completion message.

02:15–02:38 · See a boundary enforced. The last test attempts one harmless write outside the approved workspace. Look for the permission-denied result from the operating system. The useful task can finish inside its allowed directory, while this outside write is rejected. This demonstrates the file boundary exercised in this example. It does not stand in for testing every possible policy or agent.

02:38–03:00 · Remove the example resources. When the task is finished, delete the named sandbox and check that it is absent from the list. Remove the example's model connection and any temporary tunnels as well. Sandbox deletion and terminating the AWS instance are separate steps. The walkthrough keeps the resource names scoped so cleanup affects only this demonstration.

03:00–03:26 · Repeat the workflow. You can follow the included guide to repeat this workflow with your approved image and model provider. Keep the model connection private, give the agent only the tools it needs, and inspect the outputs and the policy results. Software Sushi packages the AWS environment; OpenShell applies the runtime boundaries. The guide, commands, captions, and source recording are included with this video.

The recorded test's temporary instance and volume were cleaned up. The AMI remains a private preview, its Marketplace listing remains Draft, and the separate security findings still require resolution before public release.

The boundary is enforced around actions

A prompt can ask an agent to stay in a folder. OpenShell adds a runtime boundary: policy governs the files, processes, and connections the workload can use. NVIDIA's architecture documentation explains how the gateway, trusted supervisor, and sandbox cooperate. The supervisor checks requests against policy; the sandbox contains the agent's processes.

Those rules need to match the actual workflow. Allow the required workspace and service endpoints, inspect denied actions, and review changes before broadening access. An administrator who controls the host still has responsibilities outside the agent's boundary.

A real task: read, write, verify, then test the boundary

We ran a small Python agent inside OpenShell. It used an existing Qwen model on Thor through a private connection and completed five real model responses. The model chose tools that read a sample CSV, wrote a report, independently checked its contents, and attempted one harmless write outside the allowed workspace.

Hands-on walkthrough · 1 minute 47 seconds · selected recorded commands and real results, replayed with waits shortened and the private model address replaced by a placeholder. English narration was synthesized on Thor. Watch on YouTube.
Read the hands-on transcript

00:00–00:18 · A real agent inside OpenShell. Let's run a small Python agent inside OpenShell. It will read a CSV, write a report, and check its work using a real model on Thor through a private connection. This video replays the recorded commands and results; complete logs are included.

00:18–00:35 · Set the boundaries. The agent has one CPU and a limited memory budget. Its policy allows writes in the sandbox workspace and allows Python to call only the model chat endpoint. We wait for the initial configuration to settle before starting inference.

00:35–00:51 · Let the model choose tools. Now the agent asks the model to complete the task. The model chooses the CSV reader, then supplies the report text to the writing tool. These are actual model generations and real file operations inside OpenShell.

00:51–01:08 · Verify the output. The verification tool checks the report against the original CSV. There are three tasks, totaling sixty seconds. The report is saved inside the approved workspace, and its checksum is included in the evidence.

01:08–01:28 · Show an actual denied write. The model then selects a harmless write outside the allowed workspace. The operating system returns permission denied. A separate administrator control confirms that the same user can write there without the OpenShell boundary. This is a system denial, not a model refusal.

01:28–01:47 · Clean up and repeat. Finally, we remove only the named demo sandbox and provider. The test instance, temporary disk, and private tunnels are also cleaned up. This example uses a credentialless model route, so credential injection was not exercised. The kit includes the commands, source, responses, and timings.

What the private demonstration actually established
CheckObserved resultMeaning
Model and toolsFive model responses; real CSV read, report write, and verification.A model-backed agent executed the tools inside the sandbox.
Report contentsThree sample tasks with durations totaling 60 seconds.The total describes sample data. It is not an execution-speed benchmark.
Outside writeThe operating system returned permission denied (EACCES).The attempted write was blocked by the runtime boundary.
Positive controlThe same image and user could write to that path without OpenShell.The comparison distinguishes policy enforcement from ordinary filesystem permissions.

The Python agent image was supplied separately for this demo; it is not baked into the base AMI. This run used a credentialless private model route, so it did not exercise API-key injection. It validates this particular workflow and policy, not every model, agent, or authentication path.

Choose your model and credentials separately

A hosted model service may require its own API credentials and billing account. A private local model can also be used, provided the agent can reach the correctly configured endpoint. The AMI includes neither customer keys nor model weights.

OpenShell's provider mechanism associates credentials with approved endpoints. A provider configuration and the sandbox's network policy must both permit the intended request. The inference guide explains why a service using an OpenAI-compatible protocol still needs a profile for its actual destination.

Configuring a provider is a setup step; a successful model response is a separate check. Keep credentials out of image recipes, source files, recordings, and support messages. EC2 infrastructure and model-service costs are separate from any software offer.

Questions before your first workflow

Does the model have to run on the EC2 instance?

No. In our demo, the agent ran inside the AWS sandbox while the model ran on Thor. Your deployment can use a hosted service or a private model endpoint, with the required policy and connectivity.

Can I bring my own agent?

OpenShell accepts containerized workloads. Your agent's dependencies, tools, model configuration, and policy need to be prepared for that workload. NVIDIA's pinned OpenShell source describes the minimal default image, which has no agent installed.

Can I buy the Sushi AMI today?

The candidate is a private preview. The AWS Marketplace listing is Draft, and unresolved kernel security findings must be addressed before public release. The working demo does not change that availability status.

Give the workflow a clear boundary

Start with a useful task, choose the model and agent, then define what the tools need to access. The demo shows the complete loop: a model makes decisions, tools produce a verified result, and an out-of-policy action is denied.

For help designing your own workflow, explore our LLM development services and AI pilot-to-production services, browse the AI AMIs and services we offer on AWS Marketplace, read why AI pilots stall before production, or book a discovery call to discuss a private preview.