Automating System Backups with Python

Automating System Backups with Python
Photo by Tima Miroshnichenko on Pexels

Automating System Backups with Python

Picture this: it’s 2am, a junior sysadmin fat-fingers a DROP TABLE in production, and the on-call engineer’s first question is “when was the last backup, and does it actually work?” Too often, the honest answer is “we’re not sure.” This is the article that fixes that — we’re building production-ready backup solutions that handle files, databases, compression, rotation, and cloud uploads, so “we’re not sure” is never the answer again.

Table of Contents

Why Python for Backup Automation

Python excels at backup automation because it bridges multiple domains: file operations, database connectivity, cloud APIs, scheduling, and error handling. Unlike shell scripts that become unwieldy with complex logic, Python provides robust libraries and exception handling that production environments demand.

The standard library alone gives us shutil for file operations, tarfile and zipfile for compression, subprocess for external tools, and pathlib for cross-platform path handling. Combined with third-party libraries like boto3 for AWS, paramiko for SFTP, and database-specific connectors, you have everything needed for enterprise-grade backups.

Core Backup Architecture

A production backup system requires several components working together: backup execution, compression, encryption, transfer, verification, and retention management. Your architecture should separate concerns into modular functions that can be tested independently and composed into workflows.

For IT professionals looking to deepen their understanding of Python system programming, platforms like Coursera offer specialized courses on Python for DevOps and infrastructure automation that complement what we’re building here.

Design Principles

Structure your backup scripts with these principles: idempotency (safe to run multiple times), atomicity (complete or rollback), logging (audit trail for compliance), error recovery (handle partial failures), and configurability (externalize settings from code). These aren’t nice-to-haves—they’re essential for systems running unattended at 3 AM.

Implementing File System Backups

Let’s build a comprehensive file backup system that handles incremental backups, excludes patterns, and maintains metadata. This example demonstrates production patterns you’ll use repeatedly:

import os
import shutil
import tarfile
import hashlib
import json
from pathlib import Path
from datetime import datetime
from typing import List, Dict

class FileBackupManager:
    def __init__(self, source_dirs: List[str], backup_root: str, 
                 exclude_patterns: List[str] = None):
        self.source_dirs = [Path(d) for d in source_dirs]
        self.backup_root = Path(backup_root)
        self.exclude_patterns = exclude_patterns or []
        self.manifest = {}
        
    def should_exclude(self, path: Path) -> bool:
        """Check if path matches exclusion patterns."""
        path_str = str(path)
        return any(pattern in path_str for pattern in self.exclude_patterns)
    
    def calculate_checksum(self, filepath: Path) -> str:
        """Calculate SHA256 checksum for file integrity verification."""
        sha256_hash = hashlib.sha256()
        with open(filepath, "rb") as f:
            for byte_block in iter(lambda: f.read(4096), b""):
                sha256_hash.update(byte_block)
        return sha256_hash.hexdigest()
    
    def create_backup(self, backup_name: str = None) -> Path:
        """Create compressed backup with manifest."""
        if backup_name is None:
            backup_name = f"backup_{datetime.now().strftime('%Y%m%d_%H%M%S')}"
        
        backup_dir = self.backup_root / backup_name
        backup_dir.mkdir(parents=True, exist_ok=True)
        
        archive_path = backup_dir / f"{backup_name}.tar.gz"
        
        with tarfile.open(archive_path, "w:gz") as tar:
            for source_dir in self.source_dirs:
                if not source_dir.exists():
                    print(f"Warning: Source directory {source_dir} does not exist")
                    continue
                
                for root, dirs, files in os.walk(source_dir):
                    root_path = Path(root)
                    
                    # Filter out excluded directories
                    dirs[:] = [d for d in dirs 
                              if not self.should_exclude(root_path / d)]
                    
                    for file in files:
                        file_path = root_path / file
                        
                        if self.should_exclude(file_path):
                            continue
                        
                        try:
                            # Calculate checksum before adding to archive
                            checksum = self.calculate_checksum(file_path)
                            relative_path = file_path.relative_to(source_dir.parent)
                            
                            # Store metadata
                            self.manifest[str(relative_path)] = {
                                'checksum': checksum,
                                'size': file_path.stat().st_size,
                                'modified': file_path.stat().st_mtime,
                                'backed_up': datetime.now().isoformat()
                            }
                            
                            tar.add(file_path, arcname=relative_path)
                            
                        except Exception as e:
                            print(f"Error backing up {file_path}: {e}")
        
        # Write manifest
        manifest_path = backup_dir / "manifest.json"
        with open(manifest_path, 'w') as f:
            json.dump(self.manifest, f, indent=2)
        
        print(f"Backup created: {archive_path}")
        print(f"Total files: {len(self.manifest)}")
        print(f"Archive size: {archive_path.stat().st_size / (1024**2):.2f} MB")
        
        return archive_path

# Usage example
if __name__ == "__main__":
    backup_mgr = FileBackupManager(
        source_dirs=["/var/www/html", "/etc/nginx"],
        backup_root="/backups",
        exclude_patterns=[".git", "__pycache__", ".log"]
    )
    
    backup_mgr.create_backup()

This implementation handles real-world concerns: it walks directories recursively, respects exclusion patterns, calculates checksums for integrity verification, maintains a manifest for auditing, and provides progress feedback. The manifest is crucial—it lets you verify backups without extracting them and serves as documentation for compliance.

⚠️ Common Mistake: Teams treat “backup completed” as the finish line. It isn’t. A backup you’ve never restored isn’t a backup — it’s a hope. Schedule a monthly restore drill onto a throwaway environment and verify the checksums in your manifest actually match. This is the step that separates real disaster recovery from a false sense of security.

Automating Database Backups

Database backups require different approaches than file backups. You need consistent snapshots, proper credential handling, and specialized tools. Here’s a robust implementation for PostgreSQL and MySQL that integrates with your file backup workflow:

import subprocess
import gzip
from pathlib import Path
from datetime import datetime
from typing import Optional
import os

class DatabaseBackupManager:
    def __init__(self, backup_root: str):
        self.backup_root = Path(backup_root)
        self.backup_root.mkdir(parents=True, exist_ok=True)
    
    def backup_postgresql(self, 
                         database: str,
                         host: str = "localhost",
                         port: int = 5432,
                         user: str = "postgres",
                         password: Optional[str] = None) -> Path:
        """Create compressed PostgreSQL backup using pg_dump."""
        timestamp = datetime.now().strftime('%Y%m%d_%H%M%S')
        backup_file = self.backup_root / f"postgres_{database}_{timestamp}.sql.gz"
        
        # Set password via environment variable for security
        env = os.environ.copy()
        if password:
            env['PGPASSWORD'] = password
        
        # Build pg_dump command
        cmd = [
            'pg_dump',
            '-h', host,
            '-p', str(port),
            '-U', user,
            '-F', 'p',  # Plain text format
            '-b',  # Include large objects
            '-v',  # Verbose
            database
        ]
        
        try:
            # Execute pg_dump and compress on the fly
            with gzip.open(backup_file, 'wb') as f:
                process = subprocess.Popen(
                    cmd,
                    stdout=subprocess.PIPE,
                    stderr=subprocess.PIPE,
                    env=env
                )
                
                for line in process.stdout:
                    f.write(line)
                
                process.wait()
                
                if process.returncode != 0:
                    error = process.stderr.read().decode()
                    raise Exception(f"pg_dump failed: {error}")
            
            print(f"PostgreSQL backup created: {backup_file}")
            print(f"Size: {backup_file.stat().st_size / (1024**2):.2f} MB")
            return backup_file
            
        except Exception as e:
            if backup_file.exists():
                backup_file.unlink()  # Remove partial backup
            raise Exception(f"Database backup failed: {e}")
    
    def backup_mysql(self,
                    database: str,
                    host: str = "localhost",
                    user: str = "root",
                    password: Optional[str] = None) -> Path:
        """Create compressed MySQL backup using mysqldump."""
        timestamp = datetime.now().strftime('%Y%m%d_%H%M%S')
        backup_file = self.backup_root / f"mysql_{database}_{timestamp}.sql.gz"
        
        cmd = [
            'mysqldump',
            '-h', host,
            '-u', user,
            '--single-transaction',  # Consistent snapshot for InnoDB
            '--routines',  # Include stored procedures
            '--triggers',  # Include triggers
            database
        ]
        
        if password:
            cmd.append(f'--password={password}')
        
        try:
            with gzip.open(backup_file, 'wb') as f:
                process = subprocess.Popen(
                    cmd,
                    stdout=subprocess.PIPE,
                    stderr=subprocess.PIPE
                )
                
                for line in process.stdout:
                    f.write(line)
                
                process.wait()
                
                if process.returncode != 0:
                    error = process.stderr.read().decode()
                    raise Exception(f"mysqldump failed: {error}")
            
            print(f"MySQL backup created: {backup_file}")
            return backup_file
            
        except Exception as e:
            if backup_file.exists():
                backup_file.unlink()
            raise Exception(f"Database backup failed: {e}")

# Usage example
if __name__ == "__main__":
    db_backup = DatabaseBackupManager(backup_root="/backups/databases")
    
    # Backup PostgreSQL
    db_backup.backup_postgresql(
        database="production_db",
        user="backup_user",
        password=os.getenv("POSTGRES_PASSWORD")
    )
    
    # Backup MySQL
    db_backup.backup_mysql(
        database="app_db",
        user="backup_user",
        password=os.getenv("MYSQL_PASSWORD")
    )

This database backup implementation demonstrates several critical practices: it uses environment variables for credentials instead of hardcoding them, compresses dumps on-the-fly to save disk space, validates completion before keeping files, and uses database-specific flags for consistent snapshots. The --single-transaction flag for MySQL ensures you get a consistent point-in-time snapshot without locking tables.

Compression and Retention Policies

Backups consume storage, so rotation policies are mandatory. Implement grandfather-father-son rotation or custom schemes based on your recovery objectives. Keep daily backups for a week, weekly for a month, and monthly for a year—adjust based on compliance requirements.

Python’s datetime and pathlib make rotation straightforward. Iterate backup directories, parse timestamps from filenames, and delete files outside retention windows. Always verify backups before deletion and maintain logging for audit trails.

Cloud Storage Integration

Local backups alone aren’t sufficient—you need off-site copies. AWS S3, Azure Blob Storage, and Google Cloud Storage all provide Python SDKs. The pattern is consistent: authenticate, upload with multipart for large files, set lifecycle policies for cost optimization, and enable versioning.

For boto3 with AWS S3, use the upload_file method with ExtraArgs to set storage class to Glacier for long-term retention. Implement exponential backoff for retry logic and calculate MD5 checksums for transfer verification. If you’re expanding your cloud automation skills beyond backups, DataCamp offers hands-on courses specifically focused on Python for cloud infrastructure management.

Monitoring and Logging

Production backup systems require comprehensive logging. Use Python’s logging module with rotating file handlers, structured JSON logging for parsing, and multiple severity levels. Log start times, completion times, file counts, sizes, checksums, errors, and warnings.

Integrate with monitoring systems by writing metrics to StatsD, Prometheus, or CloudWatch. Send alerts on failures using SMTP, Slack webhooks, or PagerDuty APIs. Your backup script should never fail silently—absence of errors isn’t proof of success.

Production Considerations

Before deploying backup automation to production, address these concerns: run with appropriate permissions using dedicated service accounts, implement file locking to prevent concurrent runs, handle disk space exhaustion gracefully, test restoration procedures regularly, and document recovery processes.

Schedule backups during low-activity periods using cron or systemd timers. For Windows environments, use Task Scheduler with XML configurations committed to version control. Always version control your backup scripts and configurations—your backup system is critical infrastructure.

Consider encryption for backups containing sensitive data. Use GPG for file encryption or database-native encryption. Key management is crucial—encrypted backups are useless without keys, so implement secure key storage and rotation.

Scroll to Top