Senior Platform Engineer - Cloud Architecture, Fleet Delivery, and Engineering Infrastructure
Explicitly requires vibe coding skills—directing AI coding agents through multi-step work and critically reviewing and testing their output.
About the Role
Lead the design, delivery, and operational tooling of the central cloud platform that manages software, configuration, and observability for machines deployed in production plants. Own cloud architecture, release channels to remote devices, deployment records, CI/CD, and cross-team contracts to ensure reliable OTA delivery and incident remediation.
Job Description
Role
Senior/Staff Platform Engineer responsible for the central cloud platform and engineering infrastructure that delivers, monitors, and manages software running on machines installed in customer plants. The role defines architecture and service contracts, owns release and rollback mechanisms to remote/edge systems, and provides the tooling and observability that field teams and product teams depend on.
Key Responsibilities
- Define and document platform architecture, service boundaries, data schemas, versioning and compatibility rules, and tenant isolation; author decision records and interface specifications.
- Lead the contract between edge machines and central services and provide common tooling for all product lines (edge-to-cloud protocols, compatibility gating, and versioning).
- Turn platform and infrastructure designs into operated software: decompose work, sequence across contributors, run review cadence, and be accountable for delivery dates.
- Design and operate software delivery to installed systems: release channels, staged rollouts, resumable transfers over intermittent links, rollback, and end-state verification.
- Maintain a durable, queryable source-of-truth for deployed systems (hardware, software versions, configs, change history) with reconciliation practices to detect drift.
- Ensure installed systems report health, alarms, uptime and configuration back to central views over constrained or customer-controlled networks.
- Define structured diagnostics and secure retrieval workflows for field evidence and support operations.
- Operate and own CI/CD, build agents, artifact/package hosting, code signing, test/simulation harnesses, reproducible dev environments, and release engineering for central services (including AI-agent tooling/harnesses).
- Govern infrastructure-as-code, managed databases, storage, networking, identity/secrets, containerized deployment, per-customer isolation, and per-service/tenant metrics, logs, and alerting.
Qualifications (Required)
- 7+ years professional software engineering experience with ownership of production systems.
- Demonstrated architecture ownership across teams or products with written designs/decision records/specifications.
- Experience building/operating software delivery to remote/installed machines (release channels, staged rollout, rollback, verification over intermittent/customer networks).
- Experience operating a source-of-truth for deployed state, versions, configuration, and change history with reconciliation.
- Infrastructure-as-code delivered to production (Terraform, Bicep, CloudFormation, Pulumi, or equivalent) with defensible state/drift/rollback story.
- Daily working use of containers and CI/CD; experience operating containerized services under an orchestrator or managed container service and building pipelines.
- Production use of at least one major public cloud (Azure, AWS, or GCP) covering networking, managed databases, storage, identity, TLS, and secrets.
- Strong programming in at least two languages used for services or automation; systems-level C or C++ experience, including shipping libraries for other teams.
- Linux and Git as daily tools; observability practice and instrumented systems used to find real faults.
- Experience directing AI-assisted development (vibe coding) agents through multi-step work and critically reviewing their output.
- Able to write designs and specifications and lead delivery across distributed teams.
- Eligible to work on-site in Westborough, MA.
Preferred Qualifications
- Multi-tenant SaaS isolation (per-tenant DBs, credentials, network policy).
- Kubernetes in production.
- Operated relational databases at scale (provisioning, grants, backups, migrations).
- PKI, code signing, mTLS, and certificate lifecycle management.
- OpenTelemetry or equivalent observability stack.
- Telemetry and OTA update systems for connected devices; configuration/desired-state tooling at fleet scale.
- Libraries/SDKs intended for other teams to embed and support.
- Experience in industrial, IoT, machine vision, instrumentation, or regulated/security-reviewed environments.
Work Environment
- Work from written designs and recorded decisions; all changes through pull requests and review.
- Collaborate across sites and time zones with engineering, product, and field service teams.
- Weekly delivery checkpoints and early risk reporting.
Location
On-site in Westborough, Massachusetts (full-time, permanent). Occasional travel to customer sites required.