WHY GPU MODELS NEED A REGISTRY
A model version comprises multiple artifacts: weights (5-300 GB), tokenizer, inference engine + CUDA version, quantization config, and pre/post pipelines. Without a registry, a version change in any component can cause undetected output differences.
A major API provider's incident in 2025 was caused by deploying the same weights with a newer TensorRT-LLM that changed KV cache layout, undetected for 8 hours because only weight hash was checked.
REGISTRY INFRASTRUCTURE
Three tiers: PostgreSQL for metadata, S3 for artifacts (versioned buckets), local NVMe for hot cache. MLflow extended with S3-compatible storage and parallel upload (8-16 streams) reduces 140 GB model upload from 12 min to 75 sec.
Cross-region replication: replicating 140 GB us-east-1 to eu-west-1 costs $3.50 and takes 2-4 minutes. Monthly sync cost: $100-250 for frequent deployers.
VERSIONING AND LINEAGE TRACKING
Semantic versioning: MAJOR for architecture changes, MINOR for retraining, PATCH for infrastructure changes. Each version stores training run ID, dataset version, base model, and hardware config.
Lineage enables EU AI Act compliance and cost attribution. One enterprise found a PATCH update increased memory by 15 percent, raising monthly costs $23,000. Registry allowed rapid rollback.
DEPLOYMENT PIPELINE INTEGRATION
CI/CD pipeline: upload, evaluate, approve, stage, canary at 5%, full rollout. Registry enforces constraints: prevent deploy if eval fails, block rollback to vulnerable versions, require MAJOR approval.
GPU clusters register model cache status. On deploy, registry coordinates pre-fetch to reduce cold-start from 45 to under 3 seconds.