Home / Guides / How to Back Up a Local AI Workstation

guide

How to Back Up a Local AI Workstation

Updated 2026-09-18

Protect the work that is difficult to recreate—code, datasets, fine-tunes, prompts, and configuration—without filling your backup storage with downloadable model files. This guide covers practical backup targets, desktop and server setups, recovery testing, and hardware considerations.

A local AI workstation can contain months of work, but not all of its storage has the same value. Source code, training data, fine-tuned adapters, prompt libraries, environment files, and system configuration may be difficult or impossible to recreate. Downloaded model weights, caches, and package files are often replaceable—although license restrictions, private checkpoints, or discontinued releases can change that calculation.

The best backup plan protects the work and configuration needed to rebuild the system, rather than blindly copying every large file.

A practical starting point is:

  1. Keep active work on the workstation.
  2. Maintain at least one separate local backup.
  3. Keep another copy offline or in a separate location.
  4. Test that important files can actually be restored.

What needs to be backed up?

Start by classifying data according to how difficult it would be to recreate and how quickly you would need it after a failure.

Data categoryExamplesTypical backup priority
Irreplaceable workSource code, private datasets, collected documents, annotations, research notesHighest
Expensive-to-recreate outputsFine-tunes, LoRA adapters, merged checkpoints, evaluation resultsHighest
Rebuild informationPrompts, system messages, workflows, scripts, seeds, hyperparameters, package lockfilesHigh
System configurationContainer files, compose files, service settings, driver notes, shell configurationHigh
Reproducible environmentsPackage manifests, Dockerfiles, environment files, model-serving configurationHigh
Replaceable assetsPublic model weights, package caches, temporary embeddings, generated previewsSelective
Temporary dataBuild artifacts, logs, browser caches, intermediate tensorsLow or none

Do not forget the small files

AI projects often fail to restore cleanly because the main checkpoint was saved but the surrounding information was not. Preserve:

  • Training and inference code
  • Dataset manifests and preprocessing scripts
  • Dataset versions, licenses, and source notes
  • Prompt templates and chat presets
  • Sampler settings, seeds, context lengths, and quantization choices
  • Fine-tuning configuration and hyperparameters
  • Base-model name, revision, and checksum
  • Package manifests such as requirements.txt, environment.yml, or lockfiles
  • Container definitions and startup scripts
  • Service configuration for tools such as model servers, databases, vector stores, and workflow applications
  • Custom extensions, nodes, plugins, and patches
  • Evaluation data and benchmark results
  • Documentation explaining how to rebuild the environment

A checkpoint without its base model, tokenizer, configuration, or training recipe may not be useful. A backup should preserve enough context to reproduce the result or at least continue working with it.

The main bottleneck: backup bandwidth and recovery time

The limiting factor is usually not just disk capacity. It is the time required to move and verify large datasets and checkpoint files.

A simple planning formula is:

Backup time ≈ data size / sustained transfer rate

Use the sustained rate of the complete path—not the headline speed of one component. In practice:

Effective backup rate = minimum(source read speed, network throughput, destination write speed)

For a network connection:

Bandwidth (Gbps) / 8 = theoretical GB/s

Actual throughput is lower because of protocol overhead, filesystem behavior, small files, encryption, contention, and other activity. A large sequential checkpoint may transfer efficiently, while thousands of small source files can take much longer than their total size suggests.

RPO and RTO

Two recovery targets make backup decisions clearer:

  • Recovery point objective (RPO): How much recent work can you afford to lose? An RPO of one day means losing a day of changes may be acceptable; an RPO of one hour requires more frequent backups.
  • Recovery time objective (RTO): How quickly must the workstation or service be usable again?

A personal experimentation machine may tolerate a longer RTO if code and configurations are safe. A shared inference server may need rapid recovery, snapshots, spare hardware, or a standby system.

Use the 3-2-1 rule as a baseline

A practical baseline is the 3-2-1 rule:

  • Keep at least three copies of important data.
  • Use at least two different storage types or locations.
  • Keep at least one copy offline or off-site.

This is a rule of thumb, not a guarantee. The right design depends on the value of the data, privacy requirements, internet bandwidth, and recovery needs.

For sensitive datasets, an off-site copy may need encryption before leaving the premises. For ransomware protection, an always-connected backup share is not enough by itself; use versioning, snapshots, an offline copy, or another mechanism that prevents an attacker from changing every backup.

What not to treat as a backup

Several common arrangements improve availability but do not provide a complete backup:

  • RAID: Helps a system remain available after some drive failures, but does not protect against deletion, corruption, malware, fire, or an accidental format.
  • A second internal drive: Protects against some drive failures, but not theft, power events, or a system-wide mistake.
  • A synchronized folder: Replication may also synchronize accidental deletions or corrupted files unless it provides historical versions.
  • A disk image alone: Useful for restoring an operating system, but it can include caches and replaceable model files while missing newer project data.
  • Cloud sync without version history: Convenient, but not necessarily protected from ransomware or mistaken edits.
  • A backup that has never been restored: It may contain unreadable, incomplete, or incorrectly configured data.

Practical backup configuration targets

The following targets are useful starting points. They are recommendations rather than universal requirements.

Back up by value, not by file extension

Create separate backup policies for:

  1. Critical data: Code, datasets, fine-tunes, prompts, and configuration.
  2. Useful but reproducible data: Public model weights, package caches, and downloaded tools.
  3. Temporary data: Logs, previews, generated images, and intermediate files.

Critical data should generally receive frequent versioned backups. Reproducible files can be downloaded again when needed, provided you record their source, version, and checksum.

Do not automatically exclude every model file. Back up a model when it is:

  • A private or unpublished checkpoint
  • A costly fine-tune or merged model
  • Hard to download again
  • Subject to a license that limits redistribution
  • A known-good production artifact
  • Paired with custom tokenizer or configuration files
  • Needed for a fast recovery target

Keep a manifest for large files

Instead of duplicating every replaceable model, maintain a manifest containing:

  • Repository or download location
  • Exact revision or release
  • Filename
  • Size
  • SHA-256 or other checksum
  • License and usage notes
  • Required tokenizer, adapter, or configuration
  • Quantization or conversion details

A manifest is small, searchable, and useful when rebuilding a workstation. It is not a substitute for backing up private or difficult-to-recreate weights.

Set retention deliberately

A simple version policy might keep:

  • Frequent recent versions for active project files
  • Daily versions for a shorter period
  • Weekly or monthly versions for longer-term recovery

The exact retention period should reflect how often files change and how far back you may need to recover. Versioning consumes space, so exclude caches and temporary outputs before increasing retention.

Encrypt sensitive backups

Use encryption for private datasets, credentials, proprietary prompts, and personal information. Store recovery keys separately from the backup device. If the only copy of the encryption key is on the workstation that failed, the backup may be inaccessible.

Never rely on a password stored only in a local configuration file. Include key recovery in your disaster-recovery documentation, while keeping the documentation itself protected.

Verify backups automatically and manually

A useful verification process includes:

  • Checksums for important files
  • Automatic backup job failure alerts
  • Periodic comparison of source and destination
  • A test restore of representative files
  • A complete recovery rehearsal at an interval appropriate to the system

A test restore should include more than opening a text file. Restore a project, load a representative adapter or checkpoint, install the recorded environment, and confirm that the expected service or workflow starts.

Desktop scenario: one AI workstation

A single-user desktop usually benefits from simplicity and clear separation between active data and backup data.

Example layout

  • Primary storage: Operating system, applications, active projects, and currently used model files
  • Local backup device: A separate internal or external drive dedicated to versioned backups
  • Off-site or offline copy: An encrypted removable drive rotated periodically, or an encrypted remote backup
  • Code repository: A version-control system for source code and configuration, with a separate backup of the repository if it is self-hosted

The local backup provides convenient recovery from accidental deletion or drive failure. The offline or off-site copy addresses theft, fire, ransomware, and other events that affect the whole desktop.

Desktop backup policy

Back up frequently changed work whenever the RPO requires it. For example, active code and project metadata may need multiple backups per day, while a large public model manifest may only need updating when the model inventory changes.

A practical desktop job might:

  1. Back up project directories, datasets, adapters, prompts, and configuration.
  2. Exclude caches and temporary render directories.
  3. Copy only selected model weights or production artifacts.
  4. Preserve several historical versions.
  5. Encrypt the off-site or removable copy.
  6. Send a notification when the job fails.
  7. Run a test restore periodically.

Avoid using the same physical drive for both the only active copy and the only backup copy. Also avoid leaving a removable backup drive permanently connected if it is intended to protect against ransomware.

Server scenario: shared inference or training system

A server changes the priorities. Multiple users, larger datasets, and longer-running jobs make consistent snapshots, access control, and recovery time more important.

Example layout

  • Primary storage pool: Active project data, datasets, service data, and selected model artifacts
  • Snapshot or versioning layer: Fast recovery from accidental changes or corruption
  • Separate backup target: A different storage system or backup repository
  • Off-site or offline copy: Encrypted replication, removable media, or another protected location
  • Configuration repository: Version-controlled deployment files, secrets-management instructions, and infrastructure documentation

Snapshots are useful for quick recovery, but they should not be the only copy. If the storage pool or server is lost, local snapshots may be lost with it.

Server-specific considerations

  • Coordinate backups with databases and vector stores so files are captured in a consistent state.
  • Stop, quiesce, or use the application’s supported backup method when copying active service data.
  • Keep user permissions and ownership information where the backup system supports it.
  • Record GPU drivers, container runtime versions, operating-system release, and service versions.
  • Back up deployment definitions separately from large model volumes.
  • Monitor backup duration and storage growth; a job that never finishes is not a useful backup.
  • Restrict who can delete backups or change retention policies.
  • Separate backup credentials from normal administrator credentials when possible.
  • Test recovery on spare hardware or an isolated system rather than assuming the original server will be available.

For a server with a short RTO, consider maintaining a documented rebuild procedure and a small set of known-good production artifacts. Re-downloading every model during an outage may be acceptable for a hobby system but not for a service with time-sensitive users.

Estimate backup storage before buying hardware

Estimate the space required for protected data instead of copying the entire workstation by default.

A basic capacity estimate is:

Required backup capacity ≈ (critical data × number of retained versions) + growth allowance + metadata

For versioned backups, a more realistic estimate is:

Capacity ≈ initial full copy + changed data per period × retention periods + overhead

Example:

  • Critical project data: 600 GB
  • Average new or changed data per week: 40 GB
  • One initial full copy
  • Eight weeks of changed data
  • 20% allowance for filesystem overhead and growth

Capacity ≈ 600 GB + (40 GB × 8)

Capacity ≈ 920 GB before the allowance

With the allowance:

920 GB × 1.20 ≈ 1.1 TB

This estimate covers the selected critical data, not every model and cache on the workstation. If you retain full independent copies rather than incremental or deduplicated versions, capacity requirements will be higher.

Also estimate the time needed to create and restore the initial copy. A backup destination that has enough capacity but cannot complete within the available maintenance window may require a faster connection, a different storage layout, or a narrower backup scope.

Choosing hardware for the backup path

Storage and networking decisions should follow the backup target and recovery requirements.

A local backup drive or enclosure

Prioritize:

  • Capacity for current data plus retention
  • A connection that does not create an unacceptable backup window
  • Reliable power and physical placement
  • Encryption support or an encrypted backup workflow
  • Easy replacement and monitoring
  • A clear plan for disconnecting or rotating the device

A single large drive can be simple, but it remains a single point of failure. Two rotated backup devices provide stronger protection than one permanently attached device, especially when one is stored offline.

A NAS or backup server

Prioritize:

  • Enough usable capacity after redundancy and filesystem overhead
  • Snapshot and versioning support
  • Network speed appropriate to the data volume
  • Backup software that can report failures
  • Access controls and separate backup credentials
  • A second destination or off-site replication plan

Redundant storage can reduce downtime after a drive failure, but it does not replace an independent backup. The NAS itself needs protection from deletion, ransomware, and physical loss.

Networking

For a network backup, calculate the theoretical transfer ceiling first:

Bandwidth (Gbps) / 8 = theoretical GB/s

Then plan using a lower sustained rate because real transfers include overhead and may involve many small files. If large checkpoints dominate the backup, sequential throughput matters. If source code, metadata, and millions of small files dominate, latency and file-operation performance matter more.

When planning a new workstation or server, the Build a PC tool can help evaluate the overall system and storage layout. The GPU catalog is useful when the backup plan is part of a larger AI workstation build, although the GPU itself does not protect data and should not drive backup decisions by itself.

A practical setup checklist

Data inventory

  • [ ] List project directories, datasets, adapters, prompts, and configuration files.
  • [ ] Identify private or difficult-to-recreate model files.
  • [ ] Separate critical data from caches, temporary files, and downloadable assets.
  • [ ] Record model sources, revisions, checksums, licenses, and required companion files.
  • [ ] Document the operating system, drivers, containers, packages, and service versions.

Backup design

  • [ ] Define the acceptable RPO and RTO.
  • [ ] Maintain at least three copies of important data where practical.
  • [ ] Use at least two storage types or locations.
  • [ ] Keep one copy offline or off-site.
  • [ ] Enable versioning, snapshots, or retention rather than simple mirroring.
  • [ ] Encrypt sensitive data and plan how encryption keys will be recovered.
  • [ ] Keep backup credentials separate from ordinary workstation credentials.

Operations

  • [ ] Automate scheduled backups.
  • [ ] Monitor failures, skipped files, and incomplete jobs.
  • [ ] Use checksums for important files.
  • [ ] Check that the destination has enough capacity for future growth.
  • [ ] Test restoring a project, environment, dataset sample, and representative model artifact.
  • [ ] Document the rebuild process.
  • [ ] Revisit the plan when storage, models, users, or recovery requirements change.

The efficient approach is selective protection: back up the information that preserves your work and makes the system reproducible, then handle replaceable model files according to their availability, license, and recovery cost. A smaller, verified, versioned backup is more useful than a huge untested copy of the entire workstation.

Related guides