---
layout: '@/layouts/Doc.astro'
title: 'A Small Service Mesh for My Macs and Supercomputers'
date: 2026-09-25
date-created: 2026-09-25
date-modified: 2026-09-29
description: 'How I run local dashboards, separate personal and agent knowledge vaults, shared agent memory, email, terminal history, model gateways, and persistent agent sessions across two Macs and ALCF systems -- with one health checker for the whole stack.'
---
I have accumulated a surprising number of small web services around my daily
work: a training dashboard, terminal history, separate personal and agent
knowledge vaults, an agent memory API, an email API, a model gateway, and
persistent agent sessions. None is a large application. The interesting part is
making each one available in the right places without making everything public.
This post is an inventory of that system: what each service does, how the
network paths fit together, and the commands and `launchd` patterns I use to
bring them up again. It is the implementation-level companion to [Working From
Anywhere][working-anywhere], which explains the larger remote-work design.
> [!INFO]- **TL;DR** — the pattern
>
> - Applications listen on `127.0.0.1`, not every network interface.
> - `launchd` keeps long-lived processes alive on macOS.
> - **Tailscale Serve** gives selected HTTP services private HTTPS URLs inside
> my tailnet.
> - **SSH forwards** move a port to the machine that needs it.
> - **Tailcat** is a second point-to-point path when a network blocks Tailscale's
> control plane.
> - **Cloudflare Tunnel** is reserved for one browser-facing service that must
> be reachable without joining my tailnet, and Cloudflare Access authenticates
> it.
> - A listening process is not a health check. I verify the final URL from the
> machine that will consume it.
> - One deterministic checker now verifies **14 contracts** across the local
> services, private routes, MCP connections, schedulers, and agent panes.
## The topology
The always-on MacBook, `mbph`, is the hub and the public edge for the one service
published through Cloudflare. A second MacBook, `mbpr`, is often on campus
networks and acts as a client. Aurora and CELS are remote compute environments
behind SSH bastions.
```text
mbpr (mobile Mac)
│ │
SSH forwards │ │ Tailcat forwards
▼ ▼
mbph (home Mac)
┌────────────┬──────────┴─────────┬─────────────────┐
│ │ │ │
Tailscale localhost SSH jump herdr
Serve services hosts server
│ │ │ │
▼ ▼ ▼ ▼
browsers agents Aurora / CELS Heeler / relay
```
Here is the concrete inventory. Ports are included because they make debugging
far easier than descriptions like “the notes service.”
| Service | Local backend | Reachability | What it is for |
| ----------------- | ----------------- | --------------------------------------------------------- | -------------------------------------------------------------- |
| Scrollback | `mbph:8766` | Tailscale HTTPS `:10443` | Search and read terminal history away from the terminal |
| SilverBullet | `mbph:3000` | Tailscale HTTPS `:9443` | Human-friendly editing of the shared agent-memory Markdown |
| ai-memory | `mbph:49375` | Tailscale HTTPS `:8443` and Tailcat | One persistent memory/MCP service for agents on every machine |
| AGPT dashboard | `mbph:8720` | Tailscale HTTPS `:8720` and an SSH forward to `mbpr:8720` | Monitor AuroraGPT runs from a browser |
| Personal vault | `mbph:8891` | Tailscale HTTPS `:443` | Read-only browsing and search for my Obsidian vault |
| Agent vault | loopback | Separate private Tailscale route | Git-backed history, evidence, and generated agent notes |
| llm-rosetta | `mbph:8765` | Loopback; clients reach it through the host | One OpenAI-compatible endpoint for local and ALCF models |
| Hermes gateway | local IPC | Local clients and configured messaging transports | Routes prompts, tools, MCP servers, cron, and notifications |
| AgentMail MCP | subprocess/API | Hermes MCP only | Dedicated agent inbox, drafts, send, receive, and replies |
| Herdr | local socket | Local clients and the separately authenticated relay | Persistent panes, process detection, and agent status |
| Aurora data API | Aurora `:8712` | SSH local forward to `mbph:8712` | Feed live training data into the AGPT dashboard |
| Argo / CELS | CELS HTTPS `:443` | SSH jump forward to `mbph:25939` | Reach the internal model gateway from local clients |
| Tailcat SSH | `mbph:22` | `mbpr:2222` | A backup path to SSH when Tailscale is unavailable |
| Tailcat memory | `mbph:49375` | `mbpr:8443` | The same backup path for agent memory |
| AGPT reverse view | `mbph:8720` | SSH local forward to `mbpr:8720` | Keep the dashboard at a stable localhost URL on the mobile Mac |
| Herdr relay | `mbph:8375` | Cloudflare at `relay.sf.onl` | Attach a browser to persistent terminal/agent sessions |
The table contains both applications and multiple routes to some applications.
That distinction is useful: **an application owns its local port; a transport
merely decides who can reach it.** I can replace Tailscale with an SSH forward
without changing the application.
## The common service recipe
Every local application starts life on loopback:
```bash
my-service --host 127.0.0.1 --port 9000
curl --fail http://127.0.0.1:9000/healthz
```
Binding to `127.0.0.1` means the process is not accidentally exposed on Wi-Fi,
Ethernet, or a VPN interface. I then make persistence explicit with a macOS
LaunchAgent:
```xml
Label
dev.example.my-service
ProgramArguments
/Users/me/.local/bin/my-service
RunAtLoad
KeepAlive
StandardOutPath
/Users/me/Library/Logs/my-service.log
StandardErrorPath
/Users/me/Library/Logs/my-service.err.log
```
Install it once, or restart it after a change:
```bash
label=dev.example.my-service
plist="$HOME/Library/LaunchAgents/$label.plist"
plutil -lint "$plist"
launchctl bootstrap "gui/$(id -u)" "$plist" # first install
launchctl kickstart -k "gui/$(id -u)/$label" # restart
launchctl print "gui/$(id -u)/$label" # inspect
```
`RunAtLoad` starts the job at login. `KeepAlive` restarts it if it exits. Logs
go to stable files instead of disappearing with a terminal window.
> [!WARNING]
> `KeepAlive` is not proof that a service works. It can keep restarting a
> crashing process, or keep a wedged process alive forever. Always test the
> backend and the externally consumed route separately.
## Private web apps with Tailscale Serve
[Tailscale Serve][tailscale-serve] terminates HTTPS on the tailnet hostname and
proxies to a loopback HTTP server. These are the public-safe route definitions
from `mbph`; the agent-vault route stays in private configuration:
```bash
tailscale serve --bg --https=10443 http://127.0.0.1:8766 # Scrollback
tailscale serve --bg --https=9443 http://127.0.0.1:3000 # SilverBullet
tailscale serve --bg --https=8443 http://127.0.0.1:49375 # ai-memory
tailscale serve --bg --https=8720 http://127.0.0.1:8720 # AGPT dashboard
tailscale serve --bg --https=443 http://127.0.0.1:8891 # VaultServe
tailscale serve status
```
Serve configuration persists in Tailscale, so these commands define routes;
they do not need to remain running in a shell. The backend processes still need
their own supervision.
This is **Tailscale Serve, not Funnel**. Serve is tailnet-only. Funnel would put
a service on the public internet, which is not what I want for terminal history,
notes, dashboards, or agent memory.
### Scrollback: searchable terminal history
[Scrollback][scrollback] turns captured terminal output into a small searchable
web application. This is useful when I remember seeing an error or command but
not which terminal, host, or session contained it. It is also much easier to
read long output on a phone than through a terminal multiplexer.
My wrapper allows the private Tailscale hostname while keeping the socket local:
```python
#!/usr/bin/env python3
import uvicorn
from scrollback.web.app import create_app
uvicorn.run(
create_app(allowed_hosts=["my-mac.my-tailnet.ts.net"]),
host="127.0.0.1",
port=8766,
)
```
Run the wrapper from a LaunchAgent, then add the `:10443` Serve route shown
above. The hostname allowlist matters: the proxy preserves an HTTP host that a
default local-only configuration may reject.
### SilverBullet: a human view of agent memory
[SilverBullet][silverbullet] is a Markdown knowledge base with wiki links,
search, and a pleasant browser editor. I point it at a mounted/synchronized view
of my agent-memory wiki:
```bash
mount-agent-memory
silverbullet --single \
--hostname 127.0.0.1 \
--port 3000 \
"$HOME/AgentMemory"
```
This gives humans a useful complement to programmatic retrieval: I can browse a
project's decisions, edit a durable page, or follow links across related notes.
SilverBullet in this configuration does not add application-level
authentication. Tailscale makes it private to the tailnet, **not private to one
person inside that tailnet**. On a shared tailnet I would add a restrictive
Tailscale grant/ACL for port `9443`, and preferably application authentication as
defense in depth.
### ai-memory: one memory service for every agent
[`ai-memory`][ai-memory] stores captured agent sessions plus a Git-backed
Markdown wiki. Every supported coding harness points at the same MCP endpoint,
so Claude Code, Codex, OpenCode, and other clients can retrieve the same project
history and hand work to one another.
A minimal local launch looks like:
```bash
export AI_MEMORY_AUTH_TOKEN="$(openssl rand -hex 32)"
ai-memory serve \
--transport http \
--bind 127.0.0.1:49375 \
--enable-web
```
I put the token in the LaunchAgent environment or a protected environment file,
never in a public repository. The service is exposed through Tailscale on
`:8443`; clients use the `/mcp` path and send the bearer token.
`/mcp` is a machine endpoint, **not a web page**. Opening it in a normal browser
correctly returns `401` because the browser did not send the bearer token. That
response proves only that the route reached ai-memory and that authentication
ran; it does not prove an MCP client can log in. A complete health check sends
an authenticated `initialize` request and verifies the JSON-RPC response:
```bash
curl --fail-with-body \
-H "Authorization: Bearer $AI_MEMORY_AUTH_TOKEN" \
-H 'Content-Type: application/json' \
-H 'Accept: application/json, text/event-stream' \
--data '{
"jsonrpc": "2.0",
"id": 1,
"method": "initialize",
"params": {
"protocolVersion": "2025-06-18",
"capabilities": {},
"clientInfo": {"name": "healthcheck", "version": "1.0"}
}
}' \
https://my-mac.my-tailnet.ts.net:8443/mcp
```
A healthy response is HTTP `200` with `result.serverInfo.name` equal to
`ai-memory`. The optional human web interface is a separate `/web` route and
needs browser/session authentication configured before it is useful as a page.
For a new client, the project README's installers are preferable to hand-editing
every harness:
```bash
ai-memory install-mcp --client claude-code --apply
ai-memory install-hooks --agent claude-code --apply
```
### AGPT dashboard: training state without an SSH terminal
The AGPT dashboard is a small local web app for current and historical
AuroraGPT training runs. Its backend merges cached run metadata with a live data
API forwarded from Aurora. I start the dashboard itself on `mbph`:
```bash
cd ~/agpt-dash
uv run python3 server.py --port 8720
curl --fail http://127.0.0.1:8720/
```
Tailscale Serve makes it available to my tailnet at HTTPS port `8720`. On
`mbpr`, I additionally keep the same dashboard at a predictable localhost URL:
```bash
ssh -N \
-o BatchMode=yes \
-o ControlMaster=no \
-o ControlPath=none \
-o ExitOnForwardFailure=yes \
-o ServerAliveInterval=30 \
-o ServerAliveCountMax=3 \
-L 127.0.0.1:8720:127.0.0.1:8720 \
mbph
```
That SSH command lives in a script supervised by `launchd` on `mbpr`. Disabling
SSH connection sharing is intentional: a multiplexed master can accept the
forward and detach from the process that `launchd` is supervising, leaving job
state and actual port ownership out of sync.
### VaultServe: read-only notes in a browser
VaultServe is my intentionally small, read-only Obsidian viewer. It renders
Markdown, resolves wiki links and embedded media, and searches the vault without
giving the browser a write API.
```bash
VAULT_ROOT="$HOME/Obsidian/Notes" \
VAULT_BIND=127.0.0.1 \
VAULT_PORT=8891 \
~/.hermes/vaultserve/.venv/bin/python \
~/.hermes/vaultserve/server.py
```
This is useful on devices where I want to consult notes without installing or
synchronizing the full Obsidian vault. The default Tailscale HTTPS route proxies
to it, but it remains tailnet-only.
### Two vaults, two trust boundaries
The personal vault and the agent vault use the same small Python server, but
they are separate services with separate roots, repositories, ports, and sync
rules.
The personal service on `8891` reads my Obsidian tree. Automation treats that
tree as read-only. The agent service on `8892` serves a private Git repository
containing curated memory, compiled session evidence, project histories, and a
generated index of recently modified notes. Its daily sync job stages only the
paths it owns; it cannot sweep an unrelated hand-edited history page into an
automated commit.
This split is simpler than teaching one application which pages are personal,
generated, editable, or publishable. A route now implies one source tree and one
policy.
### Hermes, AgentMail, and Herdr
The browser services are only half of the stack. `Hermes` runs the agent control
plane: model routing, MCP servers, scheduled jobs, and notifications. A dedicated
`AgentMail` inbox is connected through an MCP subprocess with a narrow tool
allowlist. I verified the full loop with one outbound message, one inbound reply,
and one threaded reply. Credentials stay in a protected environment file; the
mailbox address and provider identifiers do not belong in public configuration.
`Herdr` owns persistent terminal panes and reports process state to `Heeler`.
Current `Hermes` processes start through a Python bootstrap, so process detection
has to recognize the bootstrap signature rather than only a binary named
`hermes`. The local detector now does that without treating arbitrary Python
source containing `hermes_cli` as an agent. The three HPC panes remain attached
while `Herdr` reports each as `hermes` instead of `unknown`.
## Bringing HPC services home with SSH
Some services cannot originate on either Mac. They run behind institutional
login nodes, so SSH is the transport.
### Aurora data for the dashboard
The training-data service listens on Aurora's loopback port `8712`. A local
forward makes it look like an `mbph` service:
```bash
ssh -N \
-o ExitOnForwardFailure=yes \
-o ServerAliveInterval=30 \
-L 127.0.0.1:8712:127.0.0.1:8712 \
aurora
curl --fail http://127.0.0.1:8712/api/backbone
```
The dashboard only knows about `127.0.0.1:8712`; it does not need to know about
Aurora's bastions or network topology. This is the same indirection principle as
Tailscale Serve, in the opposite direction.
### Argo through CELS
ALCF's Argo service is HTTPS inside the CELS network. An SSH jump host carries
that endpoint back to a local TLS port:
```bash
ssh -N -f \
-o BatchMode=yes \
-o ExitOnForwardFailure=yes \
-o ServerAliveInterval=15 \
-J "$USER@logins.cels.anl.gov" \
-L 127.0.0.1:25939:apps.inside.anl.gov:443 \
"$USER@compute-01.cels.anl.gov"
```
My actual entry point is [`argo-shim`][argo-shim], which wraps the tunnel and
the authentication details:
```bash
uvx --no-cache argo-shim --host compute-01.cels.anl.gov
```
Local model gateways can now speak to `127.0.0.1:25939` while preserving the
correct upstream TLS server name. A useful acceptance test is a complete TLS
handshake, not merely seeing the SSH process:
```bash
openssl s_client \
-connect 127.0.0.1:25939 \
-servername apps.inside.anl.gov \
-brief ' all
```
On `mbpr`, one client process creates two loopback forwards:
```bash
tailcat forward '' \
2222:22 \
8443:49375
```
The first mapping makes `mbph` SSH available at `mbpr:2222`; the second makes
the ai-memory API available at `mbpr:8443`. Both commands run under LaunchAgents
with `KeepAlive` and a short throttle interval.
The node key and Tailcat address are capabilities. I keep the real values in
private configuration and use placeholders here. The key design point is that
the services above do not change: SSH still speaks SSH and ai-memory still
speaks HTTP/MCP. Only the bytes' route between the Macs changes.
## One carefully public route with Cloudflare Tunnel
Everything so far assumes the client can join my tailnet or reach an SSH path.
The exception is [Herdr][herdr]'s browser relay, which lets me attach to a
persistent terminal/agent session from an ordinary browser. That needs a public
hostname, so it gets a separate and more explicit boundary.
The relay binds to loopback on `mbph`:
```bash
HERDR_RELAY_PORT=8375 uv run ./relay/herdr_relay.py
curl -I http://127.0.0.1:8375/
```
This one detail matters more than it looks: the relay shells out to the **local**
`herdr` binary, so it can only ever report the sessions on the machine it runs
on. The relay has to live wherever the agents live. It originally ran on `mbpr`,
and once most of my agents had moved to `mbph` the public URL was faithfully
serving the wrong machine's terminals. [Moving it](#moving-the-relay-to-another-machine)
was the fix.
A named Cloudflare Tunnel publishes only that backend:
```yaml
tunnel: herdr-relay
credentials-file: /Users/me/.cloudflared/.json
ingress:
- hostname: relay.example.com
service: http://localhost:8375
- service: http_status:404
```
Create and run it with:
```bash
cloudflared tunnel login
cloudflared tunnel create herdr-relay
cloudflared tunnel route dns herdr-relay relay.example.com
cloudflared tunnel \
--config "$HOME/.cloudflared/config-herdr.yml" \
run herdr-relay
```
The final catch-all `404` prevents the tunnel from becoming an accidental
general-purpose proxy. In production I also put the hostname behind
**Cloudflare Access**; the relay retains its own token as a fallback. The tunnel
credential, access audience, and relay token live in private files and never in
the LaunchAgent plist or repository.
Two LaunchAgents supervise this path independently:
1. the Herdr relay on `127.0.0.1:8375`;
2. `cloudflared`, which maintains outbound connections to Cloudflare.
Splitting them makes failures legible. I can test the local relay first, then
the authenticated public URL, and know which half is broken.
### Moving the relay to another machine
Because the relay reports on whatever machine it runs on, "which host serves the
public URL" is a decision I expect to revisit. The useful property of a **named**
tunnel is that moving it needs no DNS change at all: `relay.example.com` is a
CNAME pointing at the tunnel's UUID, not at a host. Moving the connector is
invisible from the outside.
The move is a cutover, not a parallel run. Two copies of the relay on one LAN
collide on mDNS (both register the same hardcoded
`herdr-remote._herdr-remote._tcp.local.` name), and two connectors for one tunnel
is not a state worth reasoning about. So: stop the old host first.
Before touching anything, get any uncommitted work off the old machine. Mine had
an unpushed feature sitting in the checkout, which a migration is an excellent
way to lose:
```bash
git -C ~/projects/herdr-remote status --short
```
Install the prerequisites on the new host — `cloudflared`, `uv`, and `herdr`
itself — then clone the relay at the same commit the old host was running.
Copy the credentials and configuration. The tunnel credential and the `cert.pem`
are what let the new host claim the same tunnel; without them `cloudflared` has
no identity:
```bash
scp ~/.cloudflared/cert.pem newhost:~/.cloudflared/cert.pem
scp ~/.cloudflared/.json newhost:~/.cloudflared/
ssh newhost 'chmod 600 ~/.cloudflared/cert.pem; chmod 400 ~/.cloudflared/*.json'
```
If the two machines have different usernames — mine do — every absolute path in
the config, the env files, and the LaunchAgent plists has to be rewritten, and
then _verified_, because a plist with a bad path fails quietly:
```bash
sed 's#/Users/olduser#/Users/newuser#g' config.env > /tmp/config.env
ssh newhost 'for p in $(grep -oE "/Users/[^ ]*" ~/.config/herdr-remote/config.env); do
[ -e "$p" ] && echo "OK $p" || echo "MISS $p"
done'
```
Then cut over. Unload on the old host, confirm the port is actually released,
and only then load on the new one:
```bash
# Old host
launchctl unload ~/Library/LaunchAgents/com.herdr-remote.{relay,tunnel}.plist
lsof -nP -iTCP:8375 -sTCP:LISTEN # must be empty
# New host
launchctl load ~/Library/LaunchAgents/com.herdr-remote.{relay,tunnel}.plist
```
Verifying this is where it gets interesting, because the obvious check is
useless. `curl https://relay.example.com/` returns the same `302` before and
after the move — that is Cloudflare Access redirecting to SSO, and it would keep
returning `302` even if I had migrated nothing. The public URL cannot tell me
which machine is behind it.
Two checks actually can. The tunnel's own connector list reports the connector's
`cloudflared` version and creation time, so if the new host runs a different
build the version alone identifies it:
```bash
cloudflared tunnel info herdr-relay
```
I expect exactly one connector, created at cutover time. And then the only check
that really matters — ask the relay what it is actually serving, and confirm the
session paths belong to the new machine:
```bash
ssh newhost 'herdr pane list' | head
```
Finally, keep the old host's LaunchAgents on disk but renamed, so they do not
silently reload at next login while remaining available for rollback:
```bash
mv com.herdr-remote.relay.plist com.herdr-remote.relay.plist.migrated-YYYYMMDD
```
One log line will look alarming and is not: the relay throws a zeroconf
`NonUniqueNameException` if anything else on the LAN still advertises that mDNS
name. It runs in a daemon thread, and the relay logs `relay on :8375` and
`Polling: local` immediately afterwards. Read the lines _after_ the traceback
before concluding anything is broken.
## Starting and checking the whole stack
I used to keep this as a list of `launchctl kickstart` lines per machine, and
copy-pasted the relevant block. That was a bad habit for the reason this whole
post keeps circling: kickstarting an agent only asks `launchd` to run something.
It says nothing about whether the service came back. The list also drifted — I
added the spool drainer and the export job and never updated the snippet.
So restart commands live in one `start-services` script, while a separate
`check-services.py` verifies the steady state without changing it:
```bash
start-services # restart everything, then verify
start-services --check # verify only, change nothing
start-services --serve # also re-declare the Tailscale Serve routes
~/.hermes/scripts/check-services.py
~/.hermes/scripts/check-services.py --json
```
It picks its service list from `hostname -s`, so the same script is correct on
both machines. Each entry is a LaunchAgent label, an optional probe URL, and a
description:
```bash
read -r -d '' SERVICES_MBPH <<'EOF'
dev.saforem2.scrollback|http://127.0.0.1:8766/|Scrollback terminal history
dev.saforem2.silverbullet-agent-memory|http://127.0.0.1:3000/|SilverBullet
com.github.akitaonrails.ai-memory|http://127.0.0.1:49375/mcp|ai-memory server
dev.saforem2.vaultserve|http://127.0.0.1:8891/|VaultServe
sh.samf.tailcat-serve||Tailcat listener
com.herdr-remote.relay|http://127.0.0.1:8375/|Herdr relay
com.herdr-remote.tunnel||Cloudflare tunnel
dev.saforem2.ai-memory-hook-drain||ai-memory spool drainer
dev.saforem2.ai-memory-silverbullet-sync||ai-memory -> SilverBullet export
EOF
```
The verdict logic is the part worth stealing. `401` and `403` count as healthy:
they prove TLS, routing, and the authorization boundary all ran, and the service
is correctly refusing an unauthenticated probe. Only `000` — nothing answered at
all — is an unambiguous failure:
```bash
case "$code" in
2*|3*|401|403) printf ' %-34s ok (HTTP %s)\n' "$desc" "$code" ;;
000) printf ' %-34s NO RESPONSE\n' "$desc"; fail=$((fail+1)) ;;
*) printf ' %-34s HTTP %s\n' "$desc" "$code"; fail=$((fail+1)) ;;
esac
```
Two smaller details matter. A periodic agent such as the spool drainer shows `-`
instead of a PID between runs, which is healthy rather than missing. And the
probe URL has to be a path the service actually routes: ai-memory answers `404`
on `/` and `401` on `/mcp`, so probing `/` reports a false failure.
The checker has explicit contracts for the `Hermes` gateway, `AgentMail` MCP,
`ai-memory` MCP, `llm-rosetta`, both vaults, `SilverBullet`, the required
Tailscale Serve routes, `Herdr`, the three HPC agent panes, the `Hermes` cron
scheduler, the AGPT dashboard, and Scrollback. The local `Hermes` dashboard is
optional. A required failure exits `1`; an optional failure does not change the
exit status.
The important distinction is **protocol health versus port health**. The two MCP
checks perform real MCP connection tests. The Tailscale check parses the route
table and confirms each backend is listening. The `Herdr` check requires a
compatible live server, then verifies that the Aurora, Sunspot, and Polaris
panes are still classified as `hermes` rather than `unknown`.
The current result is **14 OK, 0 warnings, 0 failures**:
```text
OK Hermes gateway
OK AgentMail MCP
OK ai-memory MCP
OK llm-rosetta
OK personal vault
OK agent vault
OK SilverBullet
OK Tailscale
OK Herdr
OK Hermes HPC panes
OK Hermes cron
OK AGPT dashboard
OK Scrollback viewer
OK Hermes dashboard [optional]
SUMMARY ok=14 warn=0 fail=0
```
The same checker runs twice daily. Its wrapper prints nothing when every contract
is healthy, so the notification path stays quiet unless there is a warning or
failure. The JSON form carries the same verdict for other automation.
> [!WARNING]
> Restarting a service restarts everything downstream of it, and the downstream
> side may not notice. When I restarted `ai-memory` on the hub, the Tailcat
> forward on the other Mac kept listening on its local port while its far end
> was gone — `launchctl list` reported the forward as `-9`, and a request
> through it hung instead of failing. Run `start-services --check` on _both_
> machines after restarting anything shared.
## Verify paths, not processes
This stack recently produced five instructive failures:
- Tailscale Serve was healthy but proxied the AGPT URL to stale port `8726`
instead of the live dashboard on `8720`.
- VaultServe still accepted TCP connections on `8890` but closed every HTTP
request without sending a response.
- An SSH control master owned `mbpr:8720`, while the LaunchAgent that was
supposed to own the forward was not actually running.
- The Herdr relay served `relay.sf.onl` correctly from the _wrong machine_ for
weeks. Every probe passed, because the public URL returns an identical
Cloudflare Access `302` no matter which host is behind the tunnel.
- `llm-rosetta` answered correctly on IPv4 while a stale review server owned the
same port on IPv6. `127.0.0.1` returned the model list; `localhost` resolved to
the other process and returned `404`.
All five could look “up” in a process list. The acceptance checks need to cross
the same boundary as the real consumer:
```bash
# Backend first
curl --fail http://127.0.0.1:8891/healthz
# Then the tailnet route
curl --fail https://my-mac.my-tailnet.ts.net/
# Reachability/auth-boundary check only; this is not a full health check
curl -sS -o /dev/null -w '%{http_code}\n' \
https://my-mac.my-tailnet.ts.net:8443/mcp
# Test the forwarded dashboard on the consuming Mac, not the server
ssh mbpr 'curl --fail http://127.0.0.1:8720/'
# Inspect ownership as well as presence
lsof -nP -iTCP:8720 -sTCP:LISTEN
launchctl print gui/$(id -u)/sh.samf.agpt-forward-mbph
# Check both address families when localhost and 127.0.0.1 disagree
curl --fail http://127.0.0.1:8765/v1/models
curl --fail http://[::1]:8765/v1/models
```
For an authenticated API, `401` proves that TLS, routing, proxying, and the
authorization boundary ran. It does **not** prove valid credentials or protocol
initialization; use the authenticated MCP request above for that. Conversely, a
Tailscale-generated `502` means the route exists but its backend is unavailable
or failed to produce HTTP.
## Security boundaries
The useful question is not “is it encrypted?” Every path here is encrypted. The
question is **who is allowed to arrive at the application after decryption?**
| Boundary | Who can reach it | Additional control |
| ----------------- | -------------------------------------- | --------------------------------------------- |
| Loopback backend | Processes on that host | Filesystem/user permissions |
| Tailscale Serve | Devices/users admitted to the tailnet | Tailnet grants/ACLs; app auth where available |
| SSH forward | A user with an accepted SSH credential | SSH config and remote account permissions |
| Tailcat forward | A peer holding the allowed key/address | Loopback binding at the client |
| Cloudflare Tunnel | The public internet can reach the edge | Cloudflare Access plus application auth |
There are a few rules I now apply consistently:
1. **Bind local unless public is deliberate.** No application here needs
`0.0.0.0`.
2. **Treat tailnet-only as private, not secret.** Every authorized tailnet member
can potentially reach an unauthenticated Serve route unless grants say
otherwise.
3. **Keep credentials out of examples, repositories, and process arguments.**
Use protected environment files or the platform's credential store.
4. **Publish one hostname, not a network.** The Cloudflare tunnel has one ingress
rule and a terminating `404` rule.
5. **Test from the consumer.** A browser route is only healthy if a browser-side
HTTP request works; a forwarded API is only healthy if the client host can use
it.
## What this buys me
The obvious payoff is convenience: terminal history on my phone, training state
without an Aurora login shell, notes from a browser, and agent context that does
not depend on which laptop or coding harness I opened.
The larger payoff is replaceability. Each application sees a local port. Each
consumer sees a stable URL or local port. Between them I can choose Tailscale,
SSH, Tailcat, or Cloudflare according to the trust boundary and the network I am
currently on. When one transport fails, I do not have to redesign the service.
That is enough “service mesh” for two Macs and a few supercomputers: not a
cluster orchestrator, just small processes with explicit ownership, narrow
network paths, persistent supervision, and health checks that exercise the
whole route.
[working-anywhere]: https://samf.sh/posts/2026/09/22
[tailscale-serve]: https://tailscale.com/kb/1242/tailscale-serve
[scrollback]: https://github.com/saforem2/scrollback
[silverbullet]: https://silverbullet.md
[ai-memory]: https://github.com/akitaonrails/ai-memory
[argo-shim]: https://github.com/saforem2/argo-shim
[tailcat]: https://github.com/saforem2/tailcat
[herdr]: https://github.com/herdrdev/herdr