Joseandres Hinojoza

San Francisco, CA — open to relocation

Joseandres Hinojoza

Software Engineer / Site Reliability Engineer

Site Reliability Engineer and Software Engineer with 4+ years of experience building and operating distributed systems at scale. Spent 2.5 years at Google on a team operating services at 150M+ QPS, designing a monitoring system that covered 300+ clusters and analyzed 10M+ QPS of health-check traffic. Has since built observability infrastructure for LLM inference and contributed to frontier model evaluation pipelines. Focused on distributed systems, incident response, and production reliability.

2.5 yrs

Site Reliability Engineer at Google

150M+

QPS handled by team-owned services at Google

300+

Clusters covered by monitoring systems built

4+ yrs

Professional software engineering experience

Recent experience

Software Expert Engineer · Mercor

Apr 2025Present

Remote (Worldwide)

  • Validated and adapted 50+ open-source GitHub repositories to evaluate model behavior for frontier LLMs (Grok 4, Claude); resolved environment configuration issues across Python, SQL, MySQL, PostgreSQL, Docker, and Podman.
  • Designed prompt evaluation rubrics and documented model failure modes, contributing to measurable robustness improvements in AI model performance.

Software Engineer Intern · Near AI

Nov 2025Dec 2025

San Francisco, CA, United States

  • Designed and deployed a monitoring system for 5 open-source LLMs served via vLLM; implemented health probes, alerting pipelines, and real-time dashboards to ensure high availability and model reliability.
  • Configured and optimized NGINX-based load balancing across multiple inference services, improving traffic distribution and reducing tail latency.
View full experience →

Core skills

Languages

  • Python
  • C++
  • C
  • Rust
  • Java
  • JavaScript
  • Bash
  • SQL

Infrastructure & Cloud

  • Kubernetes
  • Docker
  • Podman
  • Terraform
  • Ansible
  • Borg
  • GCP
  • AWS

Observability & Reliability

  • Prometheus
  • Grafana
  • Datadog
  • PagerDuty
  • Monarch
  • SLOs / SLIs / Error Budgets
  • Chaos Engineering
  • Canary & Blue/Green Deployments
View all skills →