Site Reliability Engineer (AI Observability: 8264)
Skillhouse ·www.skillhouse.co.jp
Apply directA Global IT Service firm is seeking a Site Reliability Engineer (SRE) with a passion for building observable, reliable systems in the rapidly evolving space of large language models (LLMs) and AI agents.
In this role, you will be responsible for the reliability, performance, and operability of Analytics’ LLM-powered products by designing and maintaining observability platforms, defining SLOs, and driving incident response for AI-specific failure modes.
This position offers an opportunity to work at the intersection of AI observability, platform engineering, data engineering, and site reliability, supporting products and services utilized by millions of users worldwide.
Responsibilities:
- Design, implement, and maintain observability pipelines for LLM-based applications and AI agents, ensuring reliability, traceability, and performance at scale
- Own the reliability, scalability, and upgrade lifecycle of observability infrastructure
- Build and maintain OpenTelemetry Collector pipelines to ingest, enrich, and fan out telemetry from LLM applications to observability backends and downstream analytics systems
- Instrument and manage tracing, logging, and metrics collection for AI/ML workloads using platforms such as Langfuse, LangChain, LangSmith, and Arize AI
- Collaborate with Security to implement data masking, payload redaction, retention policies, and RBAC for traces containing PII or confidential data
- Contribute to capacity planning and TCO modeling for observability infrastructure at varying trace volumes (1x, 10x, 100x growth scenarios)
- Work with Data Platform to design reliable, idempotent export pipelines delivering trace, evaluation, cost, and prompt metadata to internal analytics systems
- Contribute to runbooks, architectural decision records (ADRs), and internal documentation for observability standards
Required skills:
- Hands-on experience deploying and operating at least one of: Langfuse, LangSmith, Arize AX, or Phoenix in a self-hosted environment
- Understanding of LLM telemetry concepts: traces, spans, observations, token/cost tracking, prompt versioning, and evaluation workflows
- Strong working knowledge of Prometheus, Grafana, and remote-write pipelines (e.g. Cortex/Thanos)
- Experience designing or operating OpenTelemetry Collector pipelines (receivers, processors, exporters)
- Familiarity with structured logging, log routing (Cloud Logging → OpenSearch/ELK), and trace correlation
- Hands-on experience with Helm, Terraform, and GitOps workflows for infrastructure lifecycle management
- Experience writing and managing observability-as-code: dashboards as JSON/Terraform, alert rules, and SLO definitions.
- Comfortable with Python for tooling, instrumentation, and automation scripting.
- Solid understanding of GCP services: GKE, GCS, Cloud SQL, Secret Manager, Workload Identity, Cloud Logging
- Experience implementing RBAC, SSO/OIDC (Okta or equivalent), and audit logging in platform tooling
- Familiarity with PII redaction, payload masking, and data retention enforcement in observability pipelines
Why should you apply:
- This is a long-term opportunity with a chance to become a permanent employee
- You will be working with international team members
- Free breakfast, lunch and dinner at the cafeteria
Company Details:
A global company with a strong presence in multiple business areas. It has achieved sustained growth both domestically and internationally, including in the U.S. and Europe. The company boasts a diverse and international environment and is committed to equal opportunity, offering a wealth of career opportunities. Due to the diverse nature of our business, we handle a wide range of technologies! You can also choose the environment you are most comfortable with, such as Windows/Mac! Meals in the company cafeteria are also free. Our chefs are always coming up with new menu items, so you can enjoy your meal without getting bored!
Working hours: 9:00 - 17:30 (Mon-Fri)
Working Style: Hybrid (4 days in office, 1 day work from home)
Holidays: Saturday, Sunday, and National Holidays, Year-end and New Year Holidays, Paid Holidays, Other Special Holidays
Services/Benefits: Social insurance, DC Pension Plan, Transportation Fee, Skillhouse University, Test payback system, and more
Frequently asked questions
Who is hiring for the Site Reliability Engineer (AI Observability: 8264) role?
Skillhouse is hiring for the Site Reliability Engineer (AI Observability: 8264) position, a Shazamme client. Apply directly on the employer's career site.
Where is the Site Reliability Engineer (AI Observability: 8264) job located?
The Site Reliability Engineer (AI Observability: 8264) role with Skillhouse is based in Setagaya-ku, JP.
What does the Site Reliability Engineer (AI Observability: 8264) role pay?
Skillhouse lists the Site Reliability Engineer (AI Observability: 8264) role at JPY 4,400–5,000 per month.
Is the Site Reliability Engineer (AI Observability: 8264) role full-time or contract?
This is a full time position at Skillhouse.
What experience level is the Site Reliability Engineer (AI Observability: 8264) role?
The Site Reliability Engineer (AI Observability: 8264) position is aimed at mid-level candidates.
How do I apply for the Site Reliability Engineer (AI Observability: 8264) role at Skillhouse?
Apply directly on Skillhouse's career page via the Apply button on this listing. ZammeJobs links straight through to the employer's ATS — no third-party form, no resume database.