Infrastructure Archive

Eight phases, written as they happened. Each entry has the diagram from that point in time, the problem that forced the change, and how it was built. Where a decision cost me something later, that is in there too. Every claim traces back to a commit in this repository.

The README once counted nine phases. I folded them into eight, grouped by the architectural shift rather than by the date. The commit history still has the finer sequence.

Phase 1: Bootstrap & GitOps polling

The first box learned to deploy itself. Slowly, and with no audit trail.

May 2026

Why

I built this by hand on one cloud VM and redeployed it over SSH. That worked twice. By the fifth change I could not tell which commit was live, so the box started polling the repository and redeploying whenever the SHA moved.

How

Nginx served the site from a Docker image I built and pushed by hand. The poller watched the repo and swapped containers on a delta. Hardening came next: HSTS, a rebuilt base image to clear CVEs, and Terraform, so the VM and its network were written down instead of remembered.

  • Deploy: a polling loop plus Docker and nginx on one VM.
  • Hardening: HSTS, a slimmer base image, a firewall that allowed only what the deploy needed.
  • IaC: Terraform took over the VM and the network.
Developer to GitHub to the box: the VM polls for a SHA delta and redeploys; Terraform declares it.

Phase 2: Observability & IaC serverless

Watch it while it runs, then write the whole stack down.

June 2026

Why

The sidecar shell scripts told me nothing until something broke. I wanted numbers while the box was healthy, and I wanted the stack in files rather than in my head. There was a second question too: could any of this run serverless?

How

An nginx stub_status feed, a small polling daemon and a systemd unit gave me live numbers, with a switch that injected failures so I could see what one looked like. The resume build moved into CI as LaTeX. Then I rebuilt the deployment on Terraform and GitHub Actions: a Go visitor tracker on Cloud Run, state in Firestore, and Workload Identity Federation so CI held no keys at all.

  • Observability: stub_status polling, a metrics daemon, and a deliberate failure switch.
  • Resume: LaTeX in CI, served from the container. It moved to HTML in Phase 4.
  • Serverless: Cloud Run, Firestore, keyless CI, a multi-stage Go image.
CI builds and deploys Cloud Run keylessly through WIF; the API writes to Firestore and the VM metrics daemon feeds the site.

Phase 3: No-Open-Ports Docker Compose

Serverless cost more than the problem was worth, so the site went back to one VM.

July 2026

Why

Artifact Registry, Cloud Run and Firestore added up. Not by much, but more than zero, for a site that gets a handful of visits a day. Serverless removed the VM maintenance and replaced it with a bill I could not justify. So everything moved back to one free-tier VM, and I used the move to change the posture: no open ports at all.

How

The tracker API went away. Nginx and cloudflared run as containers on an isolated Compose network, and every request arrives through a Cloudflare Tunnel that the host opens outward. Nothing listens on the public interface, so there is nothing to scan. I pinned the tunnel image, tightened the VM bootstrap and the nginx config, made the deploy wait for the host to answer, and put the repo under MIT.

  • Ingress: Cloudflare Tunnel only. No port is reachable from the internet.
  • Runtime: Docker Compose v2, a pinned tunnel image, a hardened nginx.
  • Pipeline: the deploy checks the host is up before it publishes.
Visitor to Cloudflare Edge to a tunnel into a host with no open ports; the image is pulled anonymously.

Phase 4: Edge migration to Cloudflare Pages

The site moved to the edge, which freed the VM for its next job.

August 2026

Why

A static resume does not need an origin server. Cloudflare Pages is faster, gives me headers I control, and costs nothing, and it left the VM idle. That was the point: the next phase needed the machine more than the site did. Publishing stayed manual on purpose. Rendering the PDF and shipping are two commands, and I would rather press the button than publish every push.

How

Pages serves the site from a _headers policy: HSTS, nosniff, no framing, a permissions policy, immutable caching for fonts, and noindex on preview URLs. CI renders the resume PDF from the HTML with headless Chrome. I added robots.txt, llms.txt and JSON-LD so both crawlers and language models can read the page without guessing.

  • Deploy: wrangler pages deploy site/, creating the project if it is missing.
  • Headers: HSTS, nosniff, X-Frame-Options, Referrer-Policy, Permissions-Policy.
  • No CSP, on purpose: the site takes no input and loads nothing from third parties. A strict policy broke the resume twice before I dropped it.
  • Resume: plain HTML and CSS, rendered to PDF by headless Chrome in CI.
  • Diagrams: self-contained pages, so light and dark are a variable swap. The one external request is a webfont.
CI renders the resume PDF and publishes to Pages; visitors reach the edge directly.

Phase 5: Repurposing the compute

Two free VMs were still running, so they got jobs.

August 2026

Why

With the site on Pages, two free-tier VMs sat idle. Shutting them down was the tidy choice. Instead I found two things worth running: an assistant I could reach from my phone, and a development box I could reach from any browser. Both kept the rule from Phase 3, so neither one is reachable from the internet.

How

Hermes went on the GCP box with no inbound ports at all. It long-polls Telegram, calls a DeepSeek model, and I administer it over IAP SSH. Then the Oracle A1.Flex arrived through Terraform and GitHub Actions: a ttyd terminal on loopback, Node from nvm, and guardrails I care about more than the rest. A tenancy budget plus quotas that zero every compute family except the Always Free shape. That turns "I hope I do not get a bill" into "this account cannot launch a paid machine".

  • Hermes host: outbound only, Telegram polling, IAP SSH for admin.
  • Oracle box: A1.Flex with 2 OCPU and 12 GB, a cloudflared tunnel, a browser terminal.
  • Guardrails: a budget and quotas that cap the tenancy at the free tier.
The freed compute: a Hermes AI host on GCP and an Oracle dev box with no open ports, both cost-guarded.

Phase 6: Vaultwarden host

The GCP box became the one thing I refuse to be relaxed about.

August 2026

Why

When Hermes moved to Oracle, the e2-micro was free again. A password vault is the workload where I am least willing to compromise. If it is reachable from the network, it is wrong. That made it a good fit for the pattern from Phase 3, and a bad place for anything clever.

How

Vaultwarden runs as one container on loopback :8000, with a cloudflared container beside it and vault.sreeramkr.com as the only way in. Secrets sit in a root-owned 0600 .env that CI writes. The Argon2 admin token contains $, which Compose would happily eat, so it is escaped on the way in. The instance runs as a service account scoped to a single private bucket. Every night the SQLite database is copied there, and the oldest of ten is pruned.

  • Reachability: a deny-all firewall, IAP-only SSH, and a listener on loopback.
  • Secrets: a 0600 .env written by CI, with signups disabled.
  • Backups: a dedicated service account, one private bucket, ten snapshots kept.
The GCP box becomes a loopback-only Vaultwarden behind the tunnel, with nightly backups to a private bucket.

Phase 7: Oracle dev box + Hermes

A general-purpose dev box, with three ways in.

September 2026

Why

Pages had taken the site and the Oracle box was sitting there. I wanted a machine I could reach from any browser, run Hermes on, and use as a dev environment. So it became all three at once, which is exactly how it ended up in the state Phase 8 cleans up.

How

The box ran Ubuntu 24.04 (aarch64) on the Always Free A1.Flex shape, provisioned by Terraform and cloud-init from this repository. ttyd served a bash login shell on loopback :7681, and the tunnel pointed ssh.sreeramkr.com at it, with Cloudflare Access in front because ttyd has no credential of its own. Node and DSH came from nvm: pnpm, @deepseek-ai/dsh@next, the dsh-tui profile, the Archify bundle. Hermes ran from the official nousresearch/hermes-agent image, with /opt/hermes mounted at /opt/data, long-polling Telegram and keeping its gateway API on loopback :8642. There was no security group. ufw denied inbound and the OpenSSH server was removed, so nothing bound a public interface. Vaultwarden stayed on the GCP box from Phase 6. A monthly pass updated apt, nvm, DSH and the Hermes image, then rebooted.

  • Ingress: an outbound tunnel and nothing else, with Access in front of the terminal.
  • Host: A1.Flex, 2 OCPU and 12 GB, no NSG, OpenSSH removed.
  • Workloads: ttyd, DSH and Hermes in Docker, all on loopback.
  • Fleet: the GCP box still runs Vaultwarden from Phase 6.
The dev box: Cloudflare Tunnel to loopback services only, ufw denying inbound, and Hermes in its own container.
The fleet as it stood: the Oracle box and the GCP Vaultwarden host, each reached only through its own outbound tunnel.

Phase 8: Consolidation, one workload and a rebuild that works (current)

One workload, one way in, and a rebuild I can trust on a bad day.

September 2026 to now

Why

By this point the box had three ways in: a browser terminal, the harness's own web UI, and an agent in a container. Each one was another thing to patch, and the agent was the only part I used. So I asked a narrower question. What is the smallest host that runs Hermes, and can it be rebuilt from nothing without losing the agent's memory?

How

Everything above the OS changed. Hermes is installed natively as a dedicated hermes user. The installer brings its own uv-managed Python, Node and ripgrep, so a monthly apt upgrade cannot break it and I am not tracking the distro's versions. Two systemd units run it: the gateway, which long-polls Telegram, and the dashboard on :9119. sshd binds loopback on :22. The dashboard listens on every interface, which is deliberate, because Hermes only switches on its username and password gate for a non-loopback bind; ufw drops inbound traffic before it arrives. Both are reachable only through the tunnel, behind Access. Docker, ttyd and the harness are gone, and CI has no shell on the box. State lives in /home/hermes/.hermes and is snapshotted every six hours to a private OCI bucket using the instance principal, so no key sits on the host. A rebuild is one workflow run with a single reset checkbox. It either empties the bucket and rebuilds for a factory-fresh agent, or keeps the snapshots and the host restores the newest one at first boot. That bucket carries prevent_destroy, so the only copy of the agent's memory cannot be planned away by accident.

  • Ingress: an outbound tunnel, with Access in front of both routes.
  • Host: A1.Flex, 2 OCPU and 12 GB, ufw denying inbound, sshd on loopback.
  • Workloads: the Hermes gateway and dashboard under systemd. No Docker, no browser terminal, no harness.
  • State: six-hourly snapshots to a private OCI bucket. A rebuild restores the newest one; a reset wipes everything.
  • Fleet: the GCP e2-micro still runs Vaultwarden.
The current fleet: one Hermes host with a snapshot-backed rebuild, and the GCP Vaultwarden host, each reached only through its own outbound tunnel.