Software Engineer (Model Inference)
coreflow · Sydney
Job description
About the role
We are looking for a Software Engineer to own and evolve our model inference stack, delivering low‑latency, high‑throughput AI services to millions of users. The role combines deep GPU‑accelerated systems work with product‑focused delivery in a fast‑moving AI entertainment company.
Key responsibilities
- Design, build, and operate a high‑throughput inference server that serves tens of millions of generations per day with sub‑200 ms latency.
- Optimize GPU utilization through batching, quantization, and custom CUDA kernels to squeeze maximum performance from our fleet.
- Productionize LoRA and other in‑house models, scaling them to serve tens of millions of users.
Required profile
- 5+ years of experience building large‑scale software, preferably in ML inference or GPU‑accelerated systems.
- Proven ability to take ideas from user research through implementation, iteration, and delivery.
- Strong drive to solve difficult problems and deliver impact at scale.
Required skills
- GPU inference (batching, quantization)
- Serving frameworks such as vLLM, TensorRT, Triton
- Experience with custom CUDA kernels
What we offer
- Top‑of‑market compensation with equity and superannuation.
- Real ownership from day one and fast‑growing scope.
- Company card for food, coffee, tools and an unlimited workspace budget.
- Visa sponsorship and relocation assistance to Sydney.
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Australia.
Salary: Software Engineer Based on 7 job offers in AustraliaSalaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 2 hours ago
Expires 1 month from now
2 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
coreflow
Sydney