About BigGeo
BigGeo is the Spatial Cloud. We help companies manage and access the world’s spatial data. Any size, any slice, any insight. Delivered in seconds.
We’re building something that hasn’t existed before: a new layer of the internet where the “where” and “when” behind every decision is instantly clear, programmable, and actionable. Our platform removes the complexity that has kept spatial data locked in silos for decades, and replaces it with speed, precision, and control.
We’re a Calgary-based company, early and moving fast, with real customers, real infrastructure, and a clear point of view on where the world is going.
Why BigGeo Exits and Why People Build Here
Most companies are spatially blind. They know what their data says, but not where or when things actually happen. That gap costs real money, creates real risk, and limits what AI can actually do in the physical world.BigGeo exists to close that gap.We’re not building another tool. We’re building the rails that connect the planet’s moving data to the systems that run the world. That’s a big problem, and it takes people who care about doing things right, not just fast.
People build here because:
- The problem is real and the category is open. We’re not competing for the middle of an existing market. We’re defining a new one. Your work shapes what the category becomes.
- Your fingerprints are on the architecture. We’re at the stage where the decisions you make today become the foundation tomorrow. What you ship matters.
- We run on clarity, not politics. We move with purpose. No bureaucratic drag, just a team that agrees on the mission and gets to work.
- You’ll grow fast because the problems are hard. Spatial data at scale is a genuinely difficult domain. If you want to be stretched, you’ll be stretched.
- We’re building for longevity. We’re not chasing hype cycles. We’re building infrastructure, the kind that compounds in value over time and earns the trust of the companies that depend on it.
The Role
BigGeo is seeking a DevOps & Site Reliability Engineer to design, automate, secure, and operate the infrastructure that powers the Spatial Cloud.
This role is responsible for building reliable cloud infrastructure, deployment systems, observability platforms, security controls, and operational tooling that enable engineering teams to deliver production software with confidence.
The role combines DevOps and Site Reliability Engineering responsibilities. You will build the systems that deliver software to production, and you will own the reliability of what runs there — service level objectives, observability, capacity planning, incident response, and the on-call practice that supports them. We combine these deliberately: engineers who build delivery systems make better reliability decisions when they also operate what they ship.
You will work closely with software engineers, platform engineers, data engineers, and product teams to ensure BigGeo’s systems remain scalable, resilient, secure, and highly available as the platform grows.
This role is ideal for someone who enjoys building infrastructure as a product, automating everything possible, and creating systems that allow engineering teams to move faster without sacrificing reliability.
What You Will Build and Own
- Cloud infrastructure supporting BigGeo production environments
- Infrastructure-as-Code frameworks and deployment pipelines
- Kubernetes clusters and container orchestration platforms
- CI/CD systems supporting engineering delivery workflows
- Monitoring, logging, alerting, and observability platforms
- Security automation and compliance controls
- Reliability engineering practices and operational standards
- Service level objectives and reliability measurement frameworks
- Incident response processes and on-call practices
- Disaster recovery and business continuity capabilities
- Cost optimization frameworks across cloud environments
- Internal developer platforms and operational tooling
Key Responsibilities
Core Responsibilities
- Design, deploy, and maintain scalable cloud infrastructure
- Build and manage Infrastructure-as-Code solutions
- Develop and optimize CI/CD pipelines for engineering teams
- Operate Kubernetes-based production environments
- Improve system reliability, performance, and fault tolerance
- Implement monitoring, observability, and alerting strategies
- Manage cloud networking, security, and access controls
- Automate operational processes and infrastructure workflows
- Lead incident response, root-cause analysis, and blameless postmortems
- Establish reliability standards, SLOs, and operational metrics
- Perform capacity planning for stateful and resource-intensive workloads
- Reduce operational toil through automation and elimination of recurring issues
- Collaborate with engineering teams to improve deployment velocity
- Evaluate and integrate AI-powered operational and automation tools
- Contribute to platform architecture decisions and infrastructure strategy
Reliability and On-Call
Reliability is treated as a core engineering responsibility at BigGeo rather than a separate function. This role carries meaningful ownership of it.
Reliability engineering
- Define and maintain service level objectives for critical services
- Build observability that makes system behavior measurable and actionable
- Perform capacity planning for both growth and failure scenarios
- Drive continuous improvement through postmortems and reliability reviews
On-call
- Production alerting is automated and routed by severity
- On-call responsibility is currently shared across the engineering team and is being formalized into a structured rotation as the team grows
- You will participate in that rotation and help define escalation paths, response expectations, and handoff practices
- New team members shadow incidents before taking primary responsibility
- On-call scheduling and supporting policies are being established as part of this work
Technology and Tools
The exact stack will evolve as the platform grows, but experience with many of the following technologies is expected:
Cloud Platforms
- AWS, Azure, or Google Cloud Platform
Infrastructure and Automation
- Terraform
- Terragrunt
- Pulumi
- Infrastructure-as-Code frameworks
- GitOps workflows
Containers and Orchestration
- Docker
- Kubernetes
- Helm
CI/CD
- GitHub Actions
- GitLab CI
- ArgoCD
- Flux
- Jenkins
Observability
- Prometheus
- Grafana
- Datadog
- OpenTelemetry
- ELK/OpenSearch
Reliability and Incident Management
- SLO and SLI frameworks
- Incident management and paging tools
- Status and escalation workflows
Security
- IAM
- Secrets management
- Cloud security tooling
- Vulnerability management
Collaboration and Operations
- Slack
- Google Workspace
- Monday
- Linear
AI Tooling
- Modern AI assistants and development tools
- AI-supported operational automation
- AI-enhanced monitoring and troubleshooting workflows
What You Bring
Required Experience
- 4+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering
- Strong experience operating cloud infrastructure in production environments
- Experience building Infrastructure-as-Code solutions
- Experience managing Kubernetes and containerized workloads
- Knowledge of CI/CD pipelines and deployment automation
- Strong Linux systems administration skills
- Experience with monitoring, logging, and observability platforms
- Understanding of networking, security, and distributed systems concepts
- Experience supporting production incident response and troubleshooting
- Experience participating in an on-call rotation for production systems
- Strong communication and cross-functional collaboration skills
Preferred Experience
- Experience supporting large-scale data platforms
- Experience with geospatial or location-based systems
- Experience building internal developer platforms
- Experience with multi-cloud environments
- Familiarity with high-volume data processing systems
- Experience implementing reliability engineering practices
- Experience defining SLOs, SLIs, and error budgets
- Experience establishing or improving on-call and incident response processes
- Knowledge of database operations and performance tuning
- Experience working within startup or high-growth technology environments
- Experience integrating AI systems into operational workflows
How This Role Contributes to the Spatial Cloud
The Spatial Cloud depends on infrastructure that is reliable, scalable, secure, and globally available.
As a DevOps & Site Reliability Engineer, you will build and operate the foundational systems that allow BigGeo to manage and access the world’s spatial data at massive scale. Your work directly impacts platform reliability, developer productivity, customer experience, and the company’s ability to deliver spatial insights in seconds.
The systems you build become critical components of the infrastructure powering the next generation of spatial computing.
Work Environment and Collaboration
BigGeo operates as a highly collaborative, AI-native startup environment.
This is an on-site role based in Calgary, Alberta. Candidates must be located in the Calgary area or willing to relocate.
Team members are expected to take ownership, move quickly, communicate clearly, and contribute beyond traditional role boundaries when needed. You will work closely with engineering, product, data, and leadership teams while helping establish the operational foundation of a category-defining company.
Success in this role requires curiosity, initiative, systems thinking, and a desire to build infrastructure that enables others to do their best work.
You will have significant influence over how BigGeo scales its platform, operations, and engineering capabilities as we continue defining the Spatial Cloud category.
Success Metrics
Success in this role means quickly developing a deep understanding of BigGeo’s production environment, taking meaningful ownership of platform reliability, and improving the systems that allow engineering teams to deploy and operate the Spatial Cloud safely at scale.
First 30 Days — Understand the Environment and Establish Operational Context
By the end of your first 30 days, you will:
- Develop a working understanding of BigGeo’s cloud infrastructure, production architecture, Kubernetes environments, networking, deployment systems, and Infrastructure-as-Code.
- Understand the current CI/CD workflows and how software moves from development through deployment into production.
- Review existing monitoring, logging, alerting, and observability coverage for critical services.
- Understand current production reliability expectations, incident response practices, escalation paths, and the developing on-call model.
- Shadow production incidents and participate in troubleshooting alongside engineering team members before taking primary on-call responsibility.
- Identify initial reliability, security, infrastructure, deployment, or operational risks and document recommended priorities.
- Establish working relationships with software, platform, data, and product teams and understand the infrastructure requirements of their workloads.
- Begin using AI-assisted tools for infrastructure analysis, troubleshooting, automation, documentation, and operational workflows.
Success at 30 days: You can explain how BigGeo’s production infrastructure operates, how software reaches production, how system health is measured, where the most important operational risks exist, and how you will begin improving reliability and developer operations.
First 60 Days — Take Ownership and Improve Reliability
By the end of your first 60 days, you will:
- Independently own meaningful components of BigGeo’s cloud infrastructure, Kubernetes environments, deployment systems, or observability stack.
- Deliver at least one material improvement to Infrastructure-as-Code, CI/CD, deployment automation, observability, or operational tooling.
- Establish or materially improve SLOs and SLIs for critical production services, creating clearer visibility into reliability and service health.
- Improve monitoring and alerting so critical production conditions are actionable while unnecessary or low-value alerts are reduced.
- Participate actively in the on-call rotation with appropriate escalation support and demonstrate effective production troubleshooting and incident response.
- Lead or materially contribute to root-cause analysis and blameless postmortems, ensuring recurring issues result in clear corrective actions.
- Identify recurring operational toil and automate at least one high-value manual infrastructure or operational workflow.
- Assess capacity requirements for critical stateful or resource-intensive workloads and identify scaling or failure risks.
- Strengthen infrastructure security, access controls, secrets management, or vulnerability management where gaps are identified.
- Use AI-assisted operational workflows to accelerate troubleshooting, infrastructure development, analysis, or automation without compromising reliability or security.
Success at 60 days: You are independently operating important parts of BigGeo’s infrastructure, contributing effectively to production response, and delivering measurable improvements to reliability, automation, observability, and engineering delivery.
First 90 Days — Strengthen the Operational Foundation of the Spatial Cloud
By the end of your first 90 days, you will:
- Own the reliability and operational health of defined production infrastructure and services, with clear SLOs, monitoring, alerting, and response expectations.
- Demonstrate measurable improvement in at least one key operational area such as deployment reliability, deployment velocity, incident frequency, recovery time, alert quality, infrastructure automation, or operational toil.
- Help establish a structured on-call practice with clear escalation paths, response expectations, handoff practices, and supporting documentation.
- Establish repeatable incident response and postmortem practices that convert production failures into system and process improvements.
- Improve infrastructure resilience through capacity planning, fault-tolerance improvements, recovery planning, or elimination of identified single points of failure.
- Advance disaster recovery and business continuity capabilities for critical infrastructure and workloads.
- Deliver reusable infrastructure automation or internal developer tooling that allows engineering teams to deploy and operate services with less manual intervention.
- Identify cloud cost or resource-efficiency opportunities and implement practical improvements without compromising reliability or performance.
- Establish clear operational documentation and runbooks for critical infrastructure, deployment, incident, and recovery workflows.
- Contribute informed recommendations to platform architecture and infrastructure strategy based on direct production experience.
- Demonstrate that infrastructure improvements are creating leverage across the engineering organization by making production systems easier to deploy, observe, troubleshoot, recover, and scale.
Success at 90 days: You are operating as a trusted owner of BigGeo’s production reliability. Engineering teams can ship with greater confidence, critical systems are more observable and resilient, operational practices are becoming repeatable, and infrastructure is better prepared to support the continued scale of the Spatial Cloud.
