Trida Network

AI Platform Engineer

Mumbai, Maharashtra, India · Engineering · Posted 1 day ago

On-siteFull-time
Sign up to apply

Already have an account? Login to Apply

AI Platform Engineer

Company: Trida Labs
Location: Mumbai, India
Work Model: Remote
Employment Type: Full-time
Experience: 3–5 Years

About Trida Labs

Trida Labs builds AI-native products that help organizations work with enterprise data, analytics, and intelligent applications.

Our products include:

  • TridaPad — An AI-native data collaboration platform for querying, analyzing, visualizing, and managing data across SQL, NoSQL, Big Data, and cloud data sources.

  • AgentSQL — An agentic AI analytics solution that enables users to interact with enterprise data using natural language, generate SQL, analyze data, create visualizations and dashboards, and derive insights.

  • TridaPad MCP Server — A Model Context Protocol server that enables AI assistants and MCP-compatible clients to interact with TridaPad and its data capabilities.

About the Role

We are looking for an AI Platform Engineer to build and operate the infrastructure and platform capabilities that support Trida Labs' AI and data products.

You will work across AI service integration, cloud infrastructure, deployment automation, customer data connectivity, security, observability, and production reliability. A key part of the role is making TridaPad reliable and secure across both hosted and self-hosted environments.

This role sits at the intersection of AI infrastructure, cloud engineering, platform engineering, security, and production operations.

The platform team owns the infrastructure, model access, deployment, observability, security, and connectivity layer, while AI Agent Engineers build agent behavior and workflows on top of these capabilities.

What You'll DoAI Platform & Model Infrastructure

  • Build reliable platform capabilities for LLMs, AI agents, analytics services, and data applications.

  • Develop reusable services and abstractions for AI model access and configuration.

  • Support model configuration, versioning, routing, fallbacks, and environment management.

  • Monitor model availability, latency, failures, usage, and resource consumption.

  • Build platform capabilities for model access, usage tracking, and AI service reliability.

  • Provide the infrastructure layer that enables AI engineers to build and operate agent workflows efficiently.

Cloud, Deployment & Infrastructure

  • Design and manage infrastructure for hosted AI and data applications.

  • Work with Render, Vercel, GCP, AWS, Azure, or equivalent platforms as required.

  • Manage compute, networking, storage, databases, and supporting services.

  • Build and maintain CI/CD pipelines for applications and platform services.

  • Containerize applications and services using Docker.

  • Automate deployment and configuration across development, staging, and production environments.

  • Improve release processes, infrastructure reliability, and operational efficiency.

Self-Hosted & Customer Data Connectivity

  • Support the deployment and operation of TridaPad in hosted and self-hosted environments.

  • Build reliable mechanisms for connecting TridaPad to customer databases and external data sources.

  • Design secure handling, encryption, rotation, and access controls for customer database credentials.

  • Support networking requirements such as controlled outbound connectivity, allowlisting, tunnels, and private database access where required.

  • Develop reliable upgrade, migration, and rollback processes for self-hosted deployments.

  • Maintain production-ready Docker images and release workflows for self-hosted installations.

  • Support secure isolation of customer data, schemas, queries, and AI context across environments.

Security & Enterprise Data Protection

  • Implement secure practices across platform infrastructure, AI services, and customer data connections.

  • Manage authentication, authorization, secrets, credentials, and service access.

  • Ensure customer credentials and sensitive configuration data are securely stored and accessed.

  • Define appropriate controls around schemas, query results, application data, and information sent to external AI services.

  • Support configuration options for customer-controlled model providers or API credentials where applicable.

  • Implement auditability for important platform and AI operations.

  • Identify and address security risks across infrastructure, deployment workflows, and customer integrations.

Observability, Reliability & Operations

  • Implement monitoring, logging, metrics, tracing, and alerting for AI and platform systems.

  • Monitor application health, service availability, latency, errors, throughput, and resource utilization.

  • Track AI service failures, model latency, token usage, and infrastructure consumption.

  • Build dashboards and operational tooling for production systems.

  • Implement health checks, retries, recovery mechanisms, and appropriate scaling strategies.

  • Investigate production incidents and perform root-cause analysis.

  • Improve platform resilience, reliability, and operational visibility across hosted and self-hosted deployments.

Performance & Cost Optimization

  • Identify performance bottlenecks across applications, infrastructure, databases, and AI services.

  • Monitor infrastructure and AI service costs.

  • Optimize compute, memory, storage, networking, model usage, and other platform resources.

  • Reduce unnecessary infrastructure and AI service consumption without compromising reliability.

  • Support capacity planning and resource optimization as workloads grow.

  • Provide usage and cost visibility to help engineering teams make informed platform decisions.

Developer Platform & Engineering Collaboration

  • Build internal tools and automation that improve engineering productivity.

  • Standardize development, testing, deployment, and operational workflows.

  • Develop reusable deployment templates, libraries, and platform components.

  • Improve local development and testing environments for AI applications.

  • Work closely with AI, backend, data, and product engineers to translate application requirements into platform solutions.

  • Support production releases, incident reviews, and continuous improvement initiatives.

  • Document platform architecture, deployment processes, security controls, and operational practices.

What We're Looking For

  • 3–5 years of experience in platform engineering, DevOps, cloud engineering, backend engineering, or a related field.

  • Strong proficiency in Python or another backend programming language.

  • Strong understanding of Linux, networking, APIs, containers, and distributed systems.

  • Hands-on experience with Docker and CI/CD.

  • Experience deploying and supporting production applications.

  • Experience with cloud or deployment platforms.

  • Experience with monitoring, logging, metrics, or distributed tracing.

  • Understanding of authentication, authorization, secrets management, and security practices.

  • Strong troubleshooting and problem-solving skills.

  • Strong communication and collaboration skills.

Preferred Qualifications

  • Experience working with LLMs, Generative AI, or AI agent applications.

  • Experience with Render, Vercel, AWS, Azure, GCP, or similar platforms.

  • Experience with Kubernetes and container orchestration.

  • Experience with Terraform or other Infrastructure as Code tools.

  • Experience with secure database connectivity and customer-managed infrastructure.

  • Experience with API gateways, message queues, caching, or asynchronous workloads.

  • Experience with AI model serving, model gateways, or LLM integrations.

  • Experience with OpenTelemetry, Prometheus, Grafana, or similar observability tools.

  • Experience with PostgreSQL, Redis, or similar infrastructure services.

  • Experience with GitHub Actions, GitLab CI, Jenkins, or similar CI/CD tools.

  • Familiarity with self-hosted or open-source software deployment and release practices.

Key Skills

Python | Platform Engineering | AI Infrastructure | Cloud | Docker | CI/CD | Linux | APIs | Distributed Systems | Infrastructure as Code | Observability | Monitoring | LLMs | Generative AI | Model Integration | AI Agents | MCP | Security | Database Connectivity | Performance Optimization

What We Offer

  • Opportunity to build infrastructure for production AI and data products.

  • Hands-on experience with AI services, enterprise data, cloud infrastructure, and self-hosted deployments.

  • Exposure to LLM infrastructure, AI agents, MCP, observability, security, and production engineering.

  • Opportunity to solve challenging problems involving customer data connectivity and reliable AI infrastructure.

  • A collaborative environment focused on building practical AI products.

Why Join Trida Labs?

At Trida Labs, you will build the platform capabilities that allow AI and data applications to operate reliably across hosted and self-hosted environments.

You will work on problems spanning AI infrastructure, customer data connectivity, deployment, security, observability, performance, and production reliability.

Equal Opportunity

Trida Labs is an equal opportunity employer. We are committed to creating an inclusive workplace where diverse perspectives are valued and everyone has the opportunity to contribute and grow.

About Trida Labs

Trida Software Labs Private Limited is an enterprise technology and software development company focused on modern data analytics and agentic AI tools.

Sign up to apply

Already have an account? Login to Apply