Zero-Downtime Docker Swarm Cluster
Employer not named by the sourceRemote
Frontier is not the employer and does not collect applications.
About this role
Linux, Amazon Web Services, Node.js, Ubuntu, Docker, Ansible, Terraform, CI/CD · I need a production-ready Docker Swarm cluster deployed on multiple Contabo VPS nodes. Each service must run with at least two replicas so traffic can shift seamlessly during updates; downtime is not acceptable.
What I expect: • A fully automated bootstrap script or Ansible/Terraform playbook that provisions the manager and worker nodes on my Contabo instances. • An opinionated but maintainable overlay network, encrypted in-flight, with service discovery and automatic rescheduling of failed containers. • Blue/green or rolling update workflow wired into a simple CI/CD step (GitHub Actions is fine) that can automatically roll back if a health-check or readiness probe fails. • Centralised logging plus real-time monitoring and alerting—Prometheus + Grafana or another open-source stack that gives me CPU, memory, container health, and custom app metrics with email/Slack alerts. • Brief but clear documentation that lets me add new nodes or services without breaking the cluster.
SSH access to freshly installed Ubuntu servers is ready; root or sudo privileges are available. I’ll test the setup by pushing an example image and performing a forced failure—if traffic keeps flowing