An AlmaLinux backup and recovery checklist helps you restore a server after a disk failure, a bad package update, ransomware, or an accidental configuration change. Recoverability depends on more than copying website files. You also need application data, database-consistent backups, system settings, repository records, recovery media, and a tested restore process.
Use this checklist for production servers, websites, applications, databases, and internal services. Record what is protected, where each copy is stored, how long it is kept, and who can restore it.
Table of Contents
1. Define the recovery priority
List each server role and decide what must return first. The order depends on how the server is used.
- Database server: Prioritize the database files or logical backups, database configuration, credentials held outside the database, and the operating system settings needed to start the database service.
- Web server: Prioritize website files, uploaded content, virtual host or web server configuration, TLS certificate records, and the database used by the site.
- Application server: Prioritize application files, deployment records, service definitions, environment settings, and the databases or queues that the application uses.
- Internal service: Prioritize service data, access settings, network configuration, and the records needed to rebuild dependent services.
Write down the required recovery order. A server may be available before the application is usable if a database, storage location, or configuration dependency is still missing.
2. Back up system and application data
Include all data that the service cannot recreate. Check the complete application data path instead of backing up only the main directory.
- Website files and user uploads
- Application data and generated files
- Database data and database-specific backup files
- Scheduled task definitions and service-related data
- SSH keys, service certificates, and other required secrets, stored with suitable access controls
- Storage mounted outside the main filesystem
Separate operating system files from application data in the backup record. This makes it clear which items can be rebuilt and which items must be restored exactly.
3. Create database-consistent backups
A database backup must represent a consistent state. Copying database files while the database is actively writing may produce an unusable or incomplete backup.
Use a database-aware backup method that follows the database system’s supported process. Record the database type, version, backup method, encryption settings, and any required restore order. Include a test restore in a separate environment or location. Confirm that the database starts, accepts a connection, and contains the expected records.
For applications with both files and a database, keep the relationship between the two parts clear. Record when each backup was created and whether the file backup and database backup can be restored together.
4. Save configurations and repository records
Configuration files are often needed to make restored data usable. Back up the relevant system and application settings, including service configuration, web server configuration, mount definitions, firewall rules, scheduled tasks, user and group records, and network settings.
Keep a record of enabled repositories, package sources, installed package versions, and important software versions. This record helps rebuild a matching environment after a failed update or disk replacement. Also record the server role, hostname, storage layout, filesystem layout, and dependencies that are not stored on the server.
5. Prepare boot and recovery media
Keep access to recovery media that can start the server when the installed system does not boot. The recovery process may be needed after a failed disk, damaged boot files, an interrupted update, or an incorrect filesystem or mount configuration.
- Confirm that recovery media can start on the server or on the replacement hardware.
- Keep the required credentials and encryption recovery information available through a controlled process.
- Record the storage layout and the location of boot, system, and application data.
- Document the steps for mounting filesystems, restoring system files, rebuilding boot information, and checking the initramfs when required.
- Store the recovery instructions separately from the server.
Test the media before it is needed. A recovery image that has not been tested may not support the required hardware, storage, or network access.
6. Set retention and off-site copies
Define how many backup versions to keep and how long to keep them. Retention should cover the time between an incident and its discovery. Keep more than one recovery point so that a damaged or encrypted backup is not the only available copy.
Store at least one copy outside the server and outside the same failure domain. Off-site storage helps when the server, local storage, or local network is unavailable. Protect backup access with separate permissions and protect sensitive backup content with encryption where required.
7. Verify backups and test restores
A completed backup job is not proof that recovery will work. Check that backup files exist, can be read, and match the expected size and date. Review failed jobs and confirm that alerts reach the responsible administrator.
Run scheduled restore tests. Restore files, configurations, and databases to a separate location, then verify permissions, ownership, service startup, application connections, and expected data. Record the restore time and any manual steps. Update the recovery procedure when the test reveals missing files or unclear dependencies.
Backups and snapshots are different
| Backup | Snapshot |
|---|---|
| A separate recovery copy intended to survive a failure of the original data. | A point-in-time view of data on the same or related storage system. |
| Can support recovery after deletion, corruption, or server loss when stored separately. | Can support quick rollback after a bad update or configuration change. |
| Requires restore testing and retention management. | Should not be treated as the only recovery copy. |
Use snapshots as a recovery aid when available, but keep independent backups for server loss, storage failure, ransomware, and long-term recovery.
8. Validate the server after recovery
After a restore or repair, confirm that the server boots normally and that all expected filesystems are mounted. Check system services, database connections, application responses, scheduled tasks, logs, permissions, and network access. Confirm that the restored data is current enough for the incident and that no required configuration was omitted.
Keep the server under observation after recovery. Review errors from the failed disk, update, or configuration change before returning the service to normal operation. Once recovery is complete, create a fresh backup and document the incident, restore point, changes made, and any improvements needed in the checklist.