Software Engineer (Site Reliability)
coreflow · Sydney
Job description
About the role
We are looking for our first dedicated Site Reliability Engineer to own reliability and core platform decisions as we scale to hundreds of millions of users.
Key responsibilities
- Own reliability and core platform decisions for a rapidly scaling AI entertainment product.
- Improve uptime and reduce recovery time objectives across critical services.
- Orchestrate and harden GPU clusters serving millions of AI generations per day.
- Implement platform‑wide observability (metrics, tracing, alerting) and enforce SLOs.
- Optimize AWS infrastructure and reduce cloud spend without sacrificing performance.
Required profile
- 5+ years operating production systems at scale.
- Strong AWS experience (infrastructure‑as‑code, high‑scale compute, Kubernetes/ECS).
- Deep observability and incident‑response expertise.
- CI/CD and deployment pipeline expertise.
- Ability to write code to fix root causes, not just symptoms.
Required skills
- AWS
- Infrastructure‑as‑Code
- Kubernetes / ECS
- CI/CD pipelines
- TypeScript
- Next.js
- React
- TailwindCSS
- tRPC
- Postgres
- Temporal
What we offer
- Top‑of‑market compensation with meaningful equity upside.
- Real ownership from day one and fast‑growing scope.
- Company card for food, coffee, tools and more.
- Daily team lunch and dinner at the office.
- Unlimited workspace budget and visa sponsorship with relocation assistance.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Australia.
Salary: Software Engineer Based on 7 job offers in AustraliaSalaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 1 hour ago
Expires 1 month from now
1 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
coreflow
Sydney