Cloud Infrastructure Engineer
Job Summary
EIRE Systems is seeking an experienced Cloud Engineer to support and enhance a large-scale private cloud environment supporting critical financial services operations. This role combines cloud infrastructure engineering, automation, observability, and platform operations across the full lifecycle of cloud services. Candidates will work in an enterprise financial services environment utilizing Linux, OpenStack, and modern automation tools in Roppongi, Tokyo.
Responsibilities
- Support and maintain large-scale OpenStack private cloud environment
- Design and implement features for infrastructure software components
- Lead integration and continuous improvement of Grafana observability ecosystem
- Troubleshoot complex issues involving Linux, virtualization, and containers
- Drive infrastructure automation and participate in on-call support rotation
Required Skills
- Minimum 5 years of experience supporting Linux systems
- Proven experience deploying and managing OpenStack at scale
- Strong experience with KVM virtualization
- Experience with Infrastructure-as-Code tools (Ansible, Terraform)
- Proficiency in Bash and Python or Go
Job Details
Position Overview
EIRE Systems is seeking an experienced Cloud Engineer to support and enhance a large-scale private cloud environment supporting critical financial services operations. This role combines cloud infrastructure engineering, automation, observability, and platform operations, providing an opportunity to work across the full lifecycle of cloud services, from design and deployment through ongoing optimization and support.
Technical Environment
• Operating Systems: Ubuntu, Red Hat Enterprise Linux, CentOS
• Cloud & Virtualization: OpenStack, KVM, Docker, LXC/LXD, MaaS
• Observability: Grafana, Prometheus/Mimir, Loki, Tempo, and IRM
• Automation & Provisioning: Ansible, Juju, Puppet, Terraform, GitLab CI/CD, and GitHub Actions
• Storage & Networking: Ceph, Cisco ACI, HAProxy, Pacemaker, Keepalived, BIND, and FreeIPA
Key Responsibilities
• Support and maintain a large-scale OpenStack private cloud environment.
• Design and implement features for infrastructure software components, managing the full lifecycle of changes from concept through production deployment.
• Lead the integration, administration, and continuous improvement of the Grafana observability ecosystem.
• Implement and maintain comprehensive monitoring across metrics, logs, and distributed tracing.
• Troubleshoot complex issues involving Linux systems, virtualization technologies, containers, and software-defined networking.
• Utilize observability data to identify bottlenecks, improve performance, and proactively resolve infrastructure issues.
• Drive infrastructure automation initiatives and support the migration of applications and services onto private cloud platforms.
• Participate in an on-call support rotation to ensure platform availability and rapid incident response.
Working Hours
09:00 - 18:00 Mon. - Fri.
*There may be occasional requests for Saturday, Sunday and Public Holiday shift work.
Job Requirements
Required Experience & Skills
• Minimum 5 years of experience supporting Linux systems, with deep knowledge of Linux internals including namespaces, cgroups, kernel tuning, and systemd.
• Proven experience deploying and managing OpenStack environments at scale.
• Strong experience with virtualization technologies, particularly KVM.
• Hands-on experience with the Grafana ecosystem or an equivalent observability platform, including metrics, logging, and distributed tracing.
• Solid understanding of networking technologies, including L2/L3 networking, VXLAN, and BGP.
• Experience working with distributed storage technologies such as Ceph.
• Professional experience with Infrastructure-as-Code tools such as Ansible, Terraform, or Juju.
• Proficiency in Bash scripting and at least one modern programming language, preferably Python or Go.
• Bachelor's degree (or higher) in Computer Science, Computer Engineering, or a related technical field.
Preferred Qualifications
• Experience deploying and supporting Kubernetes in a high-availability production environment.
• Familiarity with Canonical technologies including MaaS, Juju, and LXD.
• Experience applying Site Reliability Engineering (SRE) concepts, including SLIs, SLOs, and Error Budgets.
Language Requirements
• English Level: Business Conversation Level (TOEIC 735-860)
• Japanese Level: Daily Conversation Level