-
Notifications
You must be signed in to change notification settings - Fork 1
Operations
Navigation: Home > Pages
Version: 1.8.0-rc1
Last Updated: April 2026
Purpose: Master index for production operations documentation
This operations guide provides comprehensive documentation for deploying, managing, and troubleshooting ThemisDB in production environments with GPU acceleration.
- DevOps Engineers: Deployment and infrastructure management
- Site Reliability Engineers: Monitoring, incident response, and reliability
- Security Engineers: Security hardening and compliance
- System Administrators: Day-to-day operations and maintenance
- ML Engineers: Training and inference workload optimization
- Pre-Deployment Checklist - Verify readiness before deployment
- Deployment Guide - Step-by-step deployment instructions
- Deployment Documentation Index - All deployment-related docs
- Post-Deployment Checklist - Validation after deployment
- Kubernetes HPA Configuration - Horizontal Pod Autoscaler setup
- Scaling Runbook - Horizontal and vertical scaling procedures
- Load Balancer Integration - NGINX, AWS ALB, GCP LB, Istio, HAProxy
- Operations Documentation Index - Handbooks, runbooks, admin guides
-
Operational Runbooks - Standard operational procedures
- Upgrade Runbook - Zero-downtime upgrade procedures
- Restore Runbook - Backup restoration procedures
- Failover Runbook - Failover and recovery procedures
- Scaling Runbook - Horizontal and vertical scaling
- Monitoring Guide - Metrics, dashboards, and alerting
- Troubleshooting Guide - Common issues and solutions
- CDC Operations Runbook - Change Data Capture operational procedures
- Backup & Recovery System - Backup types, restore, PITR
- Disaster Recovery Plan - DR procedures, RTO/RPO, backup strategies
- SLA Monitoring - Service level agreements, Prometheus alerts, Grafana dashboards
- Auto-Scaling Guide - Kubernetes HPA/VPA, load balancer integration
- Security Documentation Index - All security-related docs and quick reference
- Security Hardening - Security best practices and configuration
- Compliance Checklists - SOC2, GDPR, HIPAA compliance
- Incident Response - Structured incident handling
- Operational Compliance Checklist - Monthly compliance verification
- CI/CD Documentation - Workflow architecture, release process
- CI/CD Architecture - Pipeline design and workflow details
- Release Workflows - Release automation documentation
- Maintenance Schedule - Automated maintenance tasks
- Performance Tuning - Optimization techniques and best practices
- Load Balancer Integration - Load balancer configuration and setup
- Disaster Recovery Plan - Complete DR plan with RTO/RPO targets
docs/
βββ OPERATIONS.md (this file)
βββ production/
βββ DEPLOYMENT.md # Installation and configuration
βββ DISASTER_RECOVERY_PLAN.md # Complete DR plan with RTO/RPO
βββ LOAD_BALANCER_INTEGRATION.md # Load balancer configuration
βββ PERFORMANCE_TUNING.md # Optimization guide
βββ MONITORING.md # Observability setup
βββ TROUBLESHOOTING.md # Problem resolution
βββ RUNBOOKS.md # Operational procedures
βββ SECURITY.md # Security hardening
βββ RUNBOOKS/
β βββ UPGRADE_RUNBOOK.md # Upgrade procedures
β βββ RESTORE_RUNBOOK.md # Backup restoration
β βββ FAILOVER_RUNBOOK.md # Failover & recovery
β βββ SCALING_RUNBOOK.md # Scaling operations
βββ CHECKLISTS/
β βββ pre_deployment.md # Pre-deployment validation
β βββ post_deployment.md # Post-deployment validation
β βββ incident_response.md # Incident handling
β βββ operational_compliance.md # Monthly compliance check
βββ examples/
βββ single_gpu_setup.yaml # Single GPU configuration
βββ multi_gpu_setup.yaml # Multi-GPU configuration
βββ distributed_training.yaml # Distributed training
βββ raid_configuration.yaml # High-availability storage
grafana/
βββ dashboards/
βββ sla-monitoring.json # SLA monitoring dashboard
prometheus/
βββ rules/
βββ sla-rules.yml # SLA alerting rules
helm/themisdb/
βββ templates/
βββ hpa.yaml # Horizontal Pod Autoscaler
βββ servicemonitor.yaml # Prometheus ServiceMonitor
Use Case: Development, testing, small-scale inference
Documentation:
Hardware:
- 1x RTX 3090/4090 or A100
- 64GB RAM
- 500GB NVMe SSD
Use Case: Training workloads, high-throughput inference
Documentation:
Hardware:
- 4-8x A100 or H100 GPUs
- 256GB+ RAM
- 2TB+ NVMe RAID
Use Case: Large-scale distributed training, enterprise deployments
Documentation:
Hardware:
- Multiple nodes with 4-8 GPUs each
- InfiniBand or 100 GbE networking
- Shared storage (Lustre, BeeGFS)
-
Pre-Deployment
- Complete Pre-Deployment Checklist
- Verify hardware compatibility
- Install GPU drivers and CUDA
- Configure networking and storage
-
Deployment
- Follow Deployment Guide
- Apply appropriate configuration (see examples)
- Configure security settings (Security Guide)
- Set up monitoring (Monitoring Guide)
-
Post-Deployment
- Complete Post-Deployment Checklist
- Run validation tests
- Verify monitoring and alerting
- Document deployment
Training Job:
# See detailed procedure in Runbooks
themisdb-cli job submit --config training-job.yamlInference Deployment:
# See detailed procedure in Runbooks
themisdb-cli inference deploy --model llama-2-7b# Save checkpoint
themisdb-cli checkpoint save --job-id <job-id>
# Restore from checkpoint
themisdb-cli job restore --checkpoint <checkpoint-id># Deploy LoRA adapter
themisdb-cli lora deploy --adapter custom-adapterView GPU Metrics:
-
Grafana Dashboard: http://localhost:3000
-
Prometheus: http://localhost:9090
-
GPU Status:
nvidia-smi dmon
SLA Monitoring:
- SLA Dashboard - Track availability, latency, and error budgets
- SLA Alerting Rules - Prometheus alerts for SLA breaches
- Target: 99.9% availability, P95 < 200ms, < 0.1% error rate
GPU Issues:
Performance Issues:
Training Issues:
For production incidents:
- Assess Severity (P0-Critical, P1-High, P2-Medium, P3-Low)
- Follow Incident Response Checklist: incident_response.md
- Use Troubleshooting Guide: TROUBLESHOOTING.md
- Execute Emergency Procedures: RUNBOOKS.md#emergency-procedures
| Incident | Response Guide |
|---|---|
| GPU Failure | Troubleshooting - GPU Errors |
| Out of Memory | Troubleshooting - OOM |
| Service Down | Runbooks - Emergency |
| Security Breach | Runbooks - Security Incident |
| Data Corruption | Troubleshooting - Data Corruption |
Daily:
- Monitor GPU health and utilization
- Review logs for errors
- Check disk space
- Verify backups completed
Weekly:
- Review performance metrics
- Check for GPU driver updates
- Review security logs
- Update documentation
Monthly:
- RAID scrub (if applicable)
- Security patching
- Performance tuning review
- Disaster recovery test
See Runbooks - Maintenance Windows
- Schedule maintenance window
- Follow Maintenance Procedures
- Complete Post-Deployment Checklist
- TLS 1.3 configured
- mTLS enabled for inter-node communication
- Disk encryption enabled
- Audit logging configured
- GPU access controls configured
- Key rotation automated
- Security monitoring active
Supported Standards:
- SOC 2
- GDPR
- HIPAA
-
Enable Mixed Precision: 2-3x speedup
-
Optimize Batch Size: Maximize GPU utilization
-
Enable Gradient Checkpointing: 60-80% memory savings
-
Use Flash Attention: 15-25% speedup, 30% memory reduction
| Metric | Target |
|---|---|
| GPU Utilization (Training) | >85% |
| Inference P95 Latency | <100ms |
| Training Throughput | >1000 samples/sec |
| GPU Memory Usage | <90% |
| Uptime | >99.9% |
- Check Documentation: Search this operations guide
-
Review Logs:
sudo journalctl -u themisdb -f -
Run Diagnostics:
themisdb-cli debug dump - Search Issues: https://github.com/makr-code/ThemisDB/issues
- Community Forum: https://github.com/makr-code/ThemisDB/discussions
- GitHub Issues: Bug reports and feature requests
- Discussions: Questions and community support
- Documentation: This operations guide
- Emergency: Follow on-call procedures
# Collect diagnostic information
themisdb-cli support-bundle --output /tmp/support-bundle.tar.gz
# Include in support request- Updated documentation structure and cross-links
- Added Operations, Deployment, Maintenance, and Security directory indices
- Aligned version references across all operations documents
- Added CI/CD release process links
- Initial production operations documentation
- Comprehensive deployment guides
- Performance tuning guidelines
- Security hardening procedures
- Operational runbooks
- Troubleshooting guides
- Example configurations
We welcome feedback on this documentation:
- Submit issues: https://github.com/makr-code/ThemisDB/issues
- Contribute improvements: CONTRIBUTING.md
- Discuss: https://github.com/makr-code/ThemisDB/discussions
Document Version: 1.8.1
Last Updated: May 2026
Next Review: August 2026
Quick Reference:
# Health check
themisdb-cli health
# GPU status
nvidia-smi
# Submit job
themisdb-cli job submit --config job.yaml
# Monitor job
themisdb-cli job status <job-id>
# View logs
sudo journalctl -u themisdb -f
# Backup
themisdb-cli backup create --type full
# Emergency stop
sudo systemctl stop themisdbThemisDB 1.9.0-beta Β· Home Β· Module-Index Β· GitHub Β· Issues
ThemisDB 1.9.0-beta Β· Home Β· Wiki-Index Β· Module-Index Β· FAQ Β· Quick-Reference Β· GitHub Β· Issues Β· Discussions Β· License
- Home
- Hero Articles
- All Wiki Pages
- FAQ
- Edition Comparison
- Repository README
- Changelog
- Roadmap
- Versioning
- Integration Mapping
- Overview
- Readme
- Appendix D Feature Status
- Appendix E Incident Runbooks
- Appendix F AQL Cheatsheet
- Appendix G Configuration
- Appendix H Glossary
- Appendix I Troubleshooting
- Appendix Literatur
- Chapter 00 Genesis
- Chapter 01 Introduction
- Chapter 02 Architecture
- Chapter 03 Multimodel
- Chapter 04 Installation
- Chapter 05 Relational
- Chapter 06 Graph
- Chapter 07 Document
- Chapter 08 Storage Layer
- Chapter 08 Vector
- Chapter 09 Timeseries
- Chapter 10 Enterprise
- Chapter 11 Realtime
- Chapter 12 Computervision
- Chapter 13 Fulltext
- Chapter 14 Geospatial
- Chapter 15 Analytics
- Chapter 16 Ml
- Chapter 16 Sharding
- Chapter 17 LLM Integration
- Chapter 17 Scaling
- Chapter 18 HA
- Chapter 18 Ml
- Chapter 19 Monitoring
- Chapter 19 Monitoring Observability
- Chapter 20 Backup
- Chapter 20 Performance
- Chapter 21 Auth
- Chapter 21 Performance
- Chapter 22 Clients
- Chapter 22 Encryption
- Chapter 23 Testing Qa
- Chapter 24 Ai Ethics
- Chapter 25 Devops Infrastructure
- Chapter 26 Migration Legacy
- Chapter 27 Troubleshooting
- Chapter 28 AQL Reference
- Chapter 29 Analytics Process Mining
- Chapter 30 Deployment Operations
- Chapter 31 API Protocols
- Chapter 32 API Design Rest Principles
- Chapter 32 AQL Oop Implementation
- Chapter 33 Best Practices
- Chapter 34 Query Optimization
- Chapter 35 Data Modeling Patterns
- Chapter 36 Security Hardening
- Chapter 37 Ecosystem Integration
- Chapter 38 Observability Sre
- Chapter 39 Performance Tuning Cookbook
- Chapter 40 Data Governance Compliance
- Chapter 41 Hands On Labs
- Chapter 42 Docs Assistant Usage
- Chapter MVCC Hlc
- Cover
- Cover Book
- Index
- Preface
- Test Links Example
- Batch Operations
- Best Practices
- CRUD Tutorial
- Custom Document Ingestion
- Getting Started Tutorial
- Interactive Examples
- Schema Design
- Video Tutorials
- AQL Reference
- AQL Examples
- AQL Overview
- AQL Feature Roadmap
- AQL Geospatial Guide
- AQL LLM Migration Guide
- AQL API
- AQL Grammar (EBNF)
- AQL Root Overview
- AQL Examples (root)
- API Reference
- API Module README
- OpenAPI Overview
- Client SDK Overview
- SDK Overview
- Operations
- Operations Overview
- Operations Runbook
- Operations Handbook
- ThemisCtl Admin Guide
- Pipeline E2E SOPs
- Docker Overview
- Docker Hub README
- Helm Overview
- Packaging Overview
- Operator Overview
- Security Policy
- Production Hardening Checklist
- Security Hardening Guide
- Encryption Key Management
- Access Control Framework
- Zero Trust Policy
- API Authentication & Authorization
- HSM Production Setup
- PKCS11 Integration
- DSGVO / SOC2 Checklist
- Access Model Runbooks
- Access Model Dashboard
- Maturity Automation Runbook
- Access Review Automation
- Access Model Dashboard
- Access Model Runbooks
- Rights Revocation
- Dr Checklists
- Dr Testing
- Incident Response Playbook
- Incident Response Testing
- GPU Oom Recovery
- Grammar Debugging
- Metrics Scrape Troubleshooting
- Model Swap Procedure
- Quota Tuning
- Subagent Deployment
- Logging Configuration
- Content Model
- Crypto & Keys
- Feature Flags Reference
- Modular Architecture Roadmap
- Modularization Guide
- Module Architecture Index
- PostgreSQL Wire Protocol
- Query Scheduling
- Raft Consensus Design
- Resource Pooling
- Source Directory Guide
- Unified Access Model
- E1 001 Layered Retrieval Design
- E1 002 Ann Abstraction Strategy
- E1 003 Tensor Summary Types
- E1 004 Lora Package Distinction
- E1 005 Model Switch Compatibility
- E1 006 Federated Tensor Summaries
- E2 001 Evaluation Framework Design
- E2 002 Hardware Profile Strategy
- E2 003 Query Planner Routing Model
- E2 004 Approximation Governance Rules
- E2 005 Cross Layer Fallback Confidence Policy
- E3 001 Distributed Tensor Design
- E3 002 Manifest Coordination Strategy
- E3 003 Recovery And Erasure Choice
- E3 004 Tensor Fabric Infrastructure
- Contributing
- Contributing (root)
- Code of Conduct
- Support
- Maintainers
- CTest Guide
- Build Quick Reference
- Developer Wiki Index
- Build / Test / CI
- Module Index
- Branching Strategy
- Release Strategy
- CI Policy Gates Wave C
- Disabled Stub Policy
- Docs PR Policy
- GA Promotion Sign Off
- Github Milestones Setup
- Governance Policies Phase1
- GPU Self Hosted Runner Requirements
- Hardening Phase 1 2 Summary 2026 09 23
- Maturity Claim Verification Checklist
- Maturity Evidence Registry
- Merge Gate Bot Config
- Merge Gate Status Live
- Phase 1 Closure Report
- Phase 1 Infrastructure Deployment
- Phase 1 Infrastructure Deployment Complete
- Phase 3 Baseline Capture
- Phase 3 Refinement Spec
- Phase 4 Sign Off And Closure
- Phase Closure Policy
- Phase Dependency Graph
- Phase3 Enforcement Runbook
- Plugin Submodule Rollback
- PR Version Targeting
- PR Version Targeting Backfill
- Production Ready 2026 Delivery Plan
- Publish Workflow Audit 2026 09 23
- Query Module Status
- Readme
- Release Governance
- Release Promotion Gate Policy
- Release Validation Checklist
- Root Hygiene Policy
- SBOM Approved Versions
- Security Compliance Audit Report 2026 08 10
- Security Module 5671 Evidence Summary
- Sharding P6 Residual Risk Acceptance
- Sourcecode Compliance Governance
- Src Module Documentation Compliance 2026 09 20
- Updates Development Status Sign Off
- Wave C Implementation Complete
- Wave C Implementation Plan
- Wave C Ml Exit Gate Sign Off
- Wave C Policy Gate Evidence
- Wiki Publish Tracking Guide
- Blob Storage
- Cuda
- Ethics Ai
- Exporters
- Huggingface
- Image Analysis
- Importers
- RPC
- Scraper
- Themisdb Ai Watermark Detector
- User Storage Encrypted
- Chimera Architecture
- Chimera Future
- Chimera Readme
- Chimera Roadmap
- Covina Fastapi Ingestion Architecture
- Covina Fastapi Ingestion Future
- Covina Fastapi Ingestion Roadmap
- Vcc Base Architecture
- Vcc Base Future
- Vcc Base Roadmap
- Vcc Clara Ingestion Architecture
- Vcc Clara Ingestion Future
- Vcc Clara Ingestion Roadmap
- Vcc Veritas Architecture
- Vcc Veritas Future
- Vcc Veritas Roadmap
- 01 Hello World
- 02 Todo App
- 03 Contact Manager
- 04 Inventory System
- 05 Time Series Monitor
- 06 Graph Social Network
- 07 Vector Search Documents
- 08 Dms Erp System
- 09 Iot Sensor Network
- 10 Drone Image Analysis
- 11 Blog Wiki
- 12 Expense Tracker
- 13 Recipe Manager
- 14 Ecommerce Catalog
- 15 Event Management
- 16 Kanban Board
- 17 Crm
- 18 Realtime Chat
- 19 Recommendation Engine
- 20 Smart Home
- 21 Coding Platform
- 22 AQL Diagram Tool
- 23 Traveling Salesman
- 24 Moral Philosophy Debates
- API Versioning
- Distributed Sharding
- Feedback Plugins
- Geo
- Gnn
- Image Analysis
- Legal Lora Training
- LLM
- Lora Sync
- Migration
- Nlp
- Performance
- Railway
- Replication
- Rope Visualization
- Sample Product Config
- Security
- Client SDK Overview
- Quickstart
- Sdk Enhancements
- Sdk Implementation Summary
- Test Suite Readme
- Go
- Java
- Javascript
- Php
- Python
- Ruby
- Rust
- Typescript
- 01 Grundlegende Operationen
- 02 AQL Queries
- 03 Graph Daten
- 04 Multimodell Anwendung
- 01 Quickstart Guide
- 02 AQL Referenz Kurzuebersicht
- 03 Datenmodellierung Guide
- 04 Uebungsaufgaben
- 05 Best Practices Guide
- Training Documents
- Training Overview
- 01 Einfuehrung Und Uebersicht
- 02 Datenmodelle Und Architektur
- 03 AQL Abfragesprache
- 04 Installation Und Setup
- 05 Anwendungsbeispiele
- Training Presentations
- Dependencies Readme
- Processmonitor Readme
- Themis.admintools.shared Readme
- Themis.aqlquerybuilder Readme
- Themis.aqlquerybuilder Roadmap
- Themis.auditlogviewer Readme
- Themis.auditlogviewer Roadmap
- Themis.classificationdashboard Readme
- Themis.classificationdashboard Roadmap
- Themis.compliancereports Readme
- Themis.compliancereports Roadmap
- Themis.gisviewer.controlpanel Readme
- Themis.gisviewer.controlpanel Roadmap
- Themis.impactanalysisviewer Readme
- Themis.impactanalysisviewer Roadmap
- Themis.ingestiontool Readme
- Themis.ingestiontool Roadmap
- Themis.keyrotationdashboard Readme
- Themis.keyrotationdashboard Roadmap
- Themis.piimanager Readme
- Themis.piimanager Roadmap
- Themis.retentionmanager Readme
- Themis.retentionmanager Roadmap
- Themis.sagaverifier Readme
- Themis.sagaverifier Roadmap
- Themis.usbadmintool Readme
- Themis.usbadmintool Roadmap
- Architecture Generator Readme
- CI Readme
- CI Roadmap
- Compiler Diagnostics Readme
- Compiler Diagnostics Roadmap
- Completion Readme
- Copilot Ollama Router Readme
- Copilot Ollama Router Roadmap
- Gnn Readme
- Gnn Roadmap
- Rope Visualizer Readme
- Rope Visualizer Roadmap
- Tco Calculator Readme
- Tco Calculator Roadmap
- Tests Readme
- Tests Roadmap
- Themis Config Wx Readme
- Themis Docs Builder Readme
- Wikipedia Ingestion Readme
- Ai Metadata And Provenance
- Build / Test / CI
- Governance And Roadmap
- Developer Wiki Index
- Module Direct Doxygen Check
- Module Doxygen Baseline Summary
- Module Doxygen Batch
- Module Doxygen Coverage Summary
- Module Doxygen Smoke Summary
- Modules And Apis
- Retrieval Direct Doxygen Check
- Soll Ist Gap Summary
- Wiki Delta Report