Site Reliability Engineer – OpenSearch - Remote
The Dignify Solutions, LLC
- United States, United States
- Remote
- Posted Aug 6, 2026
Job description
About the role
The remote Site Reliability Engineer – OpenSearch role is a senior (Lead/Staff) position responsible for ensuring high availability, performance, and scalability of mission‑critical cloud services by designing, building, and operating OpenSearch clusters and related infrastructure.
About the company
The Dignify Solutions, LLC is a US‑based firm offering services across accounting, human resources, education, bookkeeping, payroll, recruiting, staffing, and training.
Requirements
- 10+ years of experience as a Site Reliability Engineer specializing in OpenSearch
- Deep experience in site reliability engineering, DevOps, cloud operations, automation, observability, and distributed systems with hands‑on expertise architecting and optimizing OpenSearch clusters
- Expertise with Kubernetes, including troubleshooting, operations, management, and configuration of complex services
- Proven hands‑on experience designing, building, deploying, supporting, and maintaining OpenSearch clusters from scratch in production
- Strong experience with OpenSearch administration, cluster architecture, performance tuning, scaling, upgrades, and troubleshooting
- Experience with index design, shard and replica strategy, cluster sizing, node management, snapshot/restore, backup, and disaster recovery
- Strong understanding of distributed systems, search platforms, indexing pipelines, query optimization, and high‑availability architectures
- Expertise with Git
- Expertise with Concourse for pipeline setup, management, and troubleshooting
- Expertise with Linux, specifically SUSE and Ubuntu
- Expertise with Kafka, Zookeeper, and big‑data technologies
- Expert in developing automation for testing, deployment, scalability, and cloud service management
- Expertise building, implementing, and supporting cloud monitoring tools
- Expert knowledge of cloud computing, infrastructure operations, and databases
- Expert understanding of web services, networking, virtualization, and internet protocols
- Excellent communication and prioritization skills
- Ability to multitask and handle various projects, deadlines, and changing priorities
- Expertise with security fundamentals for SaaS multi‑tenant application systems
- Experience with AWS services including Route 53, EC2, S3, CloudWatch, DynamoDB, RDS, IAM, ACM, KMS, and VPC
- Experience deploying and operating OpenSearch in AWS‑based environments
- Experience with Cloud Foundry‑based environments
- Experience with Jenkins, Chef, and/or Terraform
- Exposure to troubleshooting IP networks and application stacks
- Experience with observability tools such as Prometheus and Grafana
- Experience with log ingestion pipelines, index lifecycle management, retention strategies, and search platform security controls