A sandbox solves a temporary problem, not an organisational one
The simplest definition of a sandbox is an isolated environment intended for work we do not want to do on our main computer or in a system shared with other teams. It might be a test of a new database version, running an unfamiliar tool, a proof of concept for a service or reproducing a bug that depends on the operating system.
A sandbox will not, however, fix a missing development process. If a project needs a permanent integration environment, automated deployments and persistent data, a single machine running for a few days will only be a workaround. The same applies when several people are meant to develop a shared service in parallel. What is needed then is an agreed topology, change control and maintenance rules, not a loose experiment.
Five signs that a separate VM is a sensible choice
The decision can rest on a few practical considerations. A sandbox VM usually makes sense when at least one of them applies:
- the test requires administrator privileges or changes to system services,
- the tool installs many dependencies and may conflict with your everyday toolchain,
- you need a specific Linux distribution, kernel version or a large amount of RAM and disk space,
- the team wants to share a working prototype without exposing a port on an employee's computer,
- after the experiment the whole environment should be deleted along with its temporary configuration.
A separate machine is also often a good way round the restrictions of a company laptop. You do not need to obtain an exception for every library or daemon. The organisation's rules on data, licences and network connections still apply, but the experiment itself does not change the local workstation.
Before you start, write down the hypothesis and the test conditions
A vague goal leads to a machine on which one tool after another is installed without answering the original question. A single sentence is often enough, for example: “We are checking whether library X can handle our data format in under five minutes and without custom extensions.” To that sentence you should add the software version, a data sample, the measurement plan and the result that will end the test.
A description prepared like this lets you choose the resources. A build test needs something different from analysing a large dataset. An experiment with a network service needs an explicit list of ports and recipients, and evaluating a tool that uses an API requires controlled credentials. The machine's specification should follow from the task, not from the wish to launch the biggest configuration possible.
Isolation has specific limits
A virtual machine separates the experiment's operating system from the user's computer, but it does not remove every risk. The code can still send traffic to the internet, download packages from external repositories or connect to company services. Before the test you therefore need to settle the rules for outbound traffic, how secrets are stored and what kind of data may be copied.
For a proof of concept it is best to use synthetic or anonymised data. API keys should have minimal scope and a validity period shorter than the machine's lifetime. If the experiment concerns a tool whose behaviour is unknown, it makes sense to limit the network to the necessary addresses and to record connection logs. Root access gives freedom, but it increases responsibility for whatever gets run.
A snapshot does not replace a description of the experiment
A snapshot lets you return quickly to a known state, for example before installing the next version. It is useful when the test covers several variants and each should start from the same point. It is not, however, documentation. Once the machine is deleted, the snapshot disappears with all its context, unless the team records the commands, versions and results somewhere permanent.
The minimum record should include the image or system identifier, a list of key packages, the test configuration, the input data, the result and the decision. For an application prototype it is worth keeping the code in a repository, and for a performance test also the load parameters and measurements. A screenshot without tool versions rarely lets you repeat the experiment later.
The end date is part of the project
A good sandbox has an expiry date. Before it arrives, the person responsible checks whether the result has made it into the repository or documentation, whether working data needs exporting and whether the credentials used have been revoked. The machine can then be shut down instead of keeping an accidental service running that nobody is watching any more.
The time-limited model also makes cost easier to assess. Instead of paying for infrastructure “just in case”, you reserve resources for the period needed to get an answer. If the result justifies further work, the next stage should have a project of its own: a repository, an owner, an integration environment and maintenance rules.
When to choose a sandbox and when another kind of environment
A single sandbox VM suits short work by one person or a small experiment with full privileges. When a dozen or so people need to carry out the same scenario, a set of identical workstations created from a single base configuration is the better option. And if the result is to run permanently and support the team's process, you need to design a maintained environment rather than extend the life of a prototype.
So the key question is not “do we need a VM?” but “what decision should this experiment inform?”. When the answer is specific, a sandbox speeds up the work. When the goal cannot be named, a separate machine merely moves the mess somewhere else.