# Sector88 > AI infrastructure for the environments the cloud cannot reach. Concatenated public content. See https://www.sector88.co/llms.txt for the curated index. --- See https://www.sector88.co for the home page overview. # AI beyond the cloud We build the infrastructure that runs AI on field hardware with strict size, weight, and power limits in environments the cloud was never designed to reach. Space. Defence. Energy. Sovereign compute. Every significant advance in computing moves from centralised systems to the point of need. Mainframes became PCs. Data centres became mobile. Cloud is becoming edge. AI follows the same path. The infrastructure for that transition does not exist yet. ## How we operate We are engineers who build for operators in the field. Our principles reflect how we think about infrastructure, deployment, and trust. ### Operator-first Every design decision starts with the person running inference in the field. If it does not make their job easier, it does not ship. ### Ship to the site We forward-deploy engineers because remote diagnostics cannot replace being in the room with the hardware. We work inside customer environments because that is where the constraints live. ### Platform, not project Seventy percent of what gets built inside a customer deployment ships back into the core product. The next customer starts where the last one finished. ## Deployed acrossborders We operate where our customers operate. Our engineering and deployment teams are positioned across time zones to support environments on every continent. ## Sector88 is the zone where AI needs to run but conventional infrastructure cannot reach. In defence and critical infrastructure, environments are divided into sectors. Zones defined by their operational constraints. We operate in the gap between what models demand and what hardware provides. ## Building toward inevitability Space computing is shipping. Orbital GPU modules are being commissioned. Defence budgets are shifting from cloud contracts to on-premise AI. Edge inference hardware is deploying faster than the software stack can keep up. The infrastructure layer to run AI on that hardware is what we build. ## We'd like to hear what you're building. Whether you are scoping a deployment or already hitting the limits of your current stack, we want to hear about it. # AI inference where you need it. One platform for AI on your own hardware. Cloud to air-gapped. ## One platform. Every environment. One install. Any hardware. Any network. We show up and make it work. ### Runs on any hardware GPU, CPU, TPU, or mixed. Runtime probes the box and configures itself for what is there. A CPU-only field server, a single Jetson at a ground station, or a rack of H100s in a SCIF. Same install, same API. ### Deploys to any environment Cloud, on-prem, edge, air-gapped, or fully disconnected. Install over a clean network or an empty one. Same Runtime. Same Hub. Same API. ### Stays on your side of the wire Your data never leaves your network. Zero egress. No metered tokens. Prompts, responses, weights, and traces stay on the hardware you installed on. ### Runtime Probes the box, picks the engine, tiers memory, serves an OpenAI-compatible API. One install. Any hardware. ### Hub One control plane for every node. Deploy, monitor, hot-swap, and roll back across your entire fleet. ### Forward-Deployed Engineers Audit, install, benchmark, harden. Our engineers embed in your team until you are live in production. ## Hardware Agnostic Any GPU, any backend, any model, anywhere. ### Hardware Platforms ### Inference Backends ### How it works under the hood Memory orchestration, engine selection, and the layers that hold it together. ### Cited engine numbers, hardware envelope Throughput across vLLM, TensorRT-LLM, and llama.cpp. Real reference deployment numbers. ### Compatibility matrix NVIDIA, AMD, Intel, Apple Silicon, ARM, CPU-only. What runs on what, with sizing per model class. ## Run it on your own hardware. Bring in our [forward-deployed engineers](/forward-deployed), or install it yourself. # Research-grade AI. Production-grade infrastructure. The platform research teams build on when the model has to run on real hardware. Constrained, sovereign, or air-gapped. Same model. Your environment. Program names and logos are property of their respective owners. ## Shrink the model. Or keep it. The conventional path makes the model smaller until it fits the hardware. We take the opposite path. Keep the model. Make the hardware reach further. ### Shrink the model Distil, quantise, prune. The model that ships is not the model that was trained. ### Keep the model. Tier the memory. We orchestrate model weights across whatever memory your system has available. The model that ships is the model that was trained. ## A systems problem. Years in the making. Memory management, kernel selection, hardware abstraction, deployment, fleet telemetry, security hardening. All of it has to work together, on every device, in every environment. ### The duct-tape stack Patched inference servers, custom CUDA kernels, bash scripts for deployment. Works on one machine until the hardware changes. ### One integrated platform Memory tiering, engine selection, fleet control, air-gapped deploy, audit, identity. Tested across Jetson, x86, GPU, CPU-only. ## Memory hierarchy.In production. Runtime probes the hardware, picks the engine, and tiers the model across whatever memory is available. GPU, system RAM, and disk. p95 latency stays predictable. The model that was trained is the model that serves. Same Runtime on a Jetson at a ground station, a research workstation, and a server rack in a SCIF. Llama-3-70B-Q4_K_M VRAM (Tier 1) 16.8 / 24 GB RAM (Tier 2) 42.3 / 64 GB SSD Cache (Tier 3) 128 / 512 GB Serving Throughput 7.8 tok/s Latency 118 ms OOM Events 0 Uptime 0s ## Three layers. One platform. ### Hub The control plane. Manage models, nodes, and deployments across every site. Audit, RBAC, telemetry, version control. ### Runtime The execution layer. Tiered memory across VRAM, RAM, NVMe. Models that should not fit run stable. GPU, CPU, or mixed. ### Deploy Single Helm chart. Air-gapped install from media. No external dependencies. Same artifact from lab to ground station to facility. ## Every node. One pane. Canary deploys, rollback on failure, hot-swap models, rotate credentials. Same model promoted across the lab, the field, and the facility. Audit trails into your SIEM. Identity through your IdP. Telemetry stays local, syncs when connectivity does. 4 3 99.9% 0 Active Deployments ground-station-08 VRAM 16.8/24 tok/s 7.8 Uptime 22d ops-center-03 VRAM 5.2/16 tok/s 24.1 Uptime 8d rig-platform-11 VRAM 6.1/8 tok/s 18.6 Uptime 45d datacenter-sg-02 VRAM -- tok/s -- Uptime 0s Activity ## Three ways in. We work with research partners in a few different shapes. Each one starts the same way: a real environment and a real problem. ### Hands-on access We give your team direct access to the platform in a real environment. You run the models, we sit alongside as you go. ### Co-funded research We co-apply on research programs where you bring the academic side and we bring the industry side. National and international funds. ### Embedded partnership Longer-term research engagements where we deploy alongside your program. Shape and terms scoped to the relationship. ## The technical shape. ### Memory orchestration Weight tiering across VRAM, RAM, and NVMe. p95 latency stays predictable under load. ### Model support Open-weight transformers, custom fine-tunes, multimodal. vLLM, Triton, llama.cpp serving paths. ### Validated hardware Jetson Orin and Orin NX, industrial edge PCs, x86 with or without GPUs, CPU-only environments. ### Air-gapped install Single Helm chart. Offline model registry. Signed bundles. Zero external dependencies. ### Audit and identity Audit trails into your SIEM. Identity through your IdP. RBAC at model and node level. ### Sovereign deployment Data stays on premises. Models stay on premises. No phone-home. No outbound calls. ## Real environments. Real constraints. ### Ground station Vision-language model on a Jetson Orin. Imagery processed at the edge. Same artifact promotes to a server rack. ### Sovereign facility LLM on classified hardware. Air-gapped install from media. No phone-home, no licence server, no outbound calls. ### University lab Multi-model benchmarking on shared, mixed-generation GPUs. Auto-detect, auto-tier, version-tracked results. ## Publishable work. Not just a deployment. Latency, throughput, and memory data across hardware. Reproducible and citable. Memory-tiered inference, edge deployment, sovereign AI. Your research, our experimental foundation. Deployment evidence, validation data, and industry letters for funding applications. ## Build with us. We work with a small number of research partners at a time. Tell us about the environment, the hardware, and the model you are trying to run. # We embed. We ship. You own it. Our engineers work inside your team until the platform runs in production. Then they hand it off. ## Not remote support. Forward-deployed. On-prem. Edge. Offshore. Classified. Air-gapped. Our engineers embed with yours, learn your hardware and constraints, then ship code that makes the Sector88 platform work in your environment. When it runs, you own it. ## Five phases. Weeks, not quarters. ### Audit Map hardware, models, and constraints. Preflight validates fit before code is written. ### Install Runtime and Hub on your infrastructure. Memory tiering tuned for your GPU, RAM, and NVMe envelope. ### Benchmark Your models, your hardware, your scorecard. Tokens/s, latency, cost vs. cloud. ### Harden OOM protection, auto-restart, audit trails, RBAC, IdP integration. Locked for your security regime. ### Live Production on your hardware, your network, your terms. You own it. ## What you walk away with. Hardware readiness audit Runtime and Hub installed Performance scorecard Production hardened Team trained and independent Improvements shipped to core ## What we learn ships back. Seventy percent of field code ships to core. Forward-deployed sits inside engineering, not sales. PRs land on main. The next customer starts where you finished. Field code ships to core PRs land on main Next customer starts ahead ## Your environment. Our engineers. Tell us what you are trying to run, and where. # Your data never leaves your network. Sector88 runs on your hardware, inside your perimeter. Zero egress by default on paid tiers. No prompts, completions, or model content ever transmitted to us. Last updated: April 2026 ## What ships today. Shipped features, not promises. This is how the platform works today. ### Zero Egress Pro and Enterprise make no outbound calls. No phone-home, no telemetry, no content collection. Licence validation is optional and runs fully offline. Your prompts, completions, model weights, and audit logs stay on the hardware you installed on. ### Air-Gapped Install Deploy from an internal registry, local mirror, or offline media. No internet dependency at any stage. Licences validated fully offline. Models loaded from local storage you control. The same install path supports classified facilities, SCIFs, and disconnected field sites. ### Encryption All API traffic encrypted with TLS 1.3. Data at rest encrypted using platform-native encryption on your infrastructure. Secrets managed through environment variables or your existing secrets manager. No keys stored by Sector88. ### Audit Logging Every administrative action is logged with user, timestamp, and action type. Prompt and response content is never captured. Logs are stored on your infrastructure, exportable in standard formats for your security team. ## Identity and access. ### Single Sign-On Hub supports SAML and OIDC. Use Okta, Azure AD, Google Workspace, or any compliant identity provider. Access delegated to your central identity management solution. ### Role-Based Access Hub roles map to your IdP groups. Administrators, operators, and viewers have distinct permission boundaries. Least-privilege defaults on every new role. ### API Security API key authentication, IP-based rate limiting, and brute-force detection on Pro and Enterprise. Prompt-free audit logs for every request. OpenAI-compatible endpoint hardened for production use. ## How we build software. ### Secure SDLC Code review required on every change. Static analysis and dependency vulnerability scanning in CI. Secure coding guidelines established and followed across the engineering team. ### Supply Chain All dependencies pinned to exact versions. Automated vulnerability scanning on every build. Software bill of materials (SBOM) available on request for enterprise customers. ### Patch Management Critical vulnerabilities patched and released within documented SLAs. Version notifications available through Hub for fleet-wide awareness. Customers control their own update schedule. ## Exactly what Community sends. Community Edition sends one anonymous heartbeat on first run and every 24 hours. The full schema is published below. Paid tiers send nothing. Set `S88_TELEMETRY=off` to disable completely. ## Compliance roadmap. Active engagements with current status. Contact us for the latest reports and documentation. ### SOC 2 Type II Controls in place, audit engagement starting. Contact us for a pre-audit summary of controls. ### ISO 27001 ISMS scoped and gap analysis underway. Targeting certification within 12 months. ### CISA Secure by Design Engineering practices aligned with the CISA Secure by Design pledge. Formal sign-on in progress. ### CSA STAR Level 1 Consensus Assessments Initiative Questionnaire (CAIQ) drafted. Self-assessment submission planned. ## Security questions. ### Does Sector88 see my prompts or completions? No. Sector88 Runtime and Hub run entirely on your infrastructure. Prompts, completions, model weights, and audit logs never leave your network. We have no mechanism to access inference content. ### Is Sector88 SOC 2 compliant? SOC 2 Type II is in progress. Controls are in place and an audit engagement is starting. [Contact us](/contact) for a pre-audit summary of our current controls. ### Can I run fully air-gapped with no outbound calls? Yes. Enterprise deployments support fully offline operation with zero outbound calls, no licence pings, and no telemetry. Deploy from an internal mirror or offline media. ### Does Sector88 encrypt data? All API traffic uses TLS 1.3. Data at rest is encrypted using your infrastructure's native encryption. We do not store or transmit any customer data on our systems. ### Can you complete a security questionnaire? Yes. We complete vendor security questionnaires, CAIQ, and SIG assessments. [Contact us](/contact) and we will return the completed questionnaire within a reasonable timeframe. ### Where can I find the telemetry schema? The complete telemetry schema is published in the Telemetry section above. Community Edition only. Paid tiers send nothing by default. ## Report a vulnerability. If you believe you have found a security issue in Sector88 Runtime, Hub, or the website, we want to hear from you. Include a description, reproduction steps, affected version, and any proof-of-concept material. ### Contact [security@sector88.co](mailto:security@sector88.co) ### Response Time We acknowledge reports within 2 business days and work with you to validate and resolve the issue. ### Coordinated Disclosure We request 90 days from initial report before public disclosure. We will keep you informed of remediation progress throughout. ### Safe Harbour Good-faith security research conducted under this policy is authorized and will not be pursued legally. We will not initiate or support legal action against researchers who comply with this disclosure policy. See also: [security.txt](/.well-known/security.txt) ## Questions about security? Our engineering team can walk through the architecture, answer questionnaires, or discuss your compliance requirements.