How I would run the overnight TechOps seat for a fleet of APEX exchanges, and a working console built for it.
Sole responder for customer exchanges while New York sleeps. Asia is in business hours and Europe is starting its day.
Debian, Proxmox, Docker, pfSense, MySQL, Ansible across cloud, bare metal and VPS. SOC 2 Type 2 change control on top.
A running monitor: four APEX gateways probed every minute with recorded uptime, events and Prometheus metrics, plus a live feed, host triage and the 9am handoff.
| Takeaway | Why it matters overnight |
|---|---|
| Alert on change against a baseline | One live gateway lists 34 of 307 pairs as Stopped. A "not Running" alert would page every night and train people to ignore pages. |
| Two vantage points before paging | A socket that fails from one network and succeeds from the edge is a path problem. Paging the matching-engine owner for it costs trust. |
| Know what to leave alone | Skipping a replica row, forcing cluster quorum or toggling CARP can turn a contained incident into a ledger or split-brain problem. |
| The handoff is a deliverable | The day shift should act on the note at 09:00 without a single follow-up question. |
White-label exchange infrastructure since 2013. APEX bundles the OMS, matching engine, gateway and asset manager; customers get a branded venue. For TechOps, every customer is a production environment to keep up.
| Area | What is public | TechOps read |
|---|---|---|
| Platform | APEX: OMS, matching engine, order routing, asset manager; custody and KYC integrations. WebSocket API with {m,i,n,o} frames. | Gateway, OMS, market data and wallets are separate failure domains. |
| Hosting | An older Microsoft case study describes Azure. The current TechOps posting lists Proxmox, Docker, pfSense and cloud, bare metal and VPS. | A mixed estate. Runbooks need a variant per hosting type. |
| Assurance | SOC 1 and SOC 2 Type 1 and Type 2 examinations completed for APEX. | Night changes need tickets and evidence, emergency ones included. |
| Customers | Named in company case studies and press: NDAX (Canada), Coinext (Brazil), Banexcoin, Wenia (Bancolombia), Green-X (Labuan), CME / TRM RMG. | Regulated operators expect clear, timestamped incident updates. |
| New surface | Perpetual futures technology (Aug 2025), a treasury platform (2026), POLYX support (Oct 2025), open hardware and wallet engineering roles. | More services and possibly hardware in the fleet within a year. |
| Coverage model | Overnight TechOps roles posted for remote and for Manila; a day-shift EST role also exists. | Follow-the-sun. The handoff quality decides whether it works. |
Crypto trades 24/7, but ticket volume follows operator business hours. The overnight window covers the full Asian working day and the European morning.
| Region | Local, during EDT | Local, during EST | What to expect |
|---|---|---|---|
| New York, Toronto | 01:00 to 09:00 | 01:00 to 09:00 | Quiet order flow; the batch jobs, backups and patch windows that commonly run overnight. |
| São Paulo, Bogotá | 02:00 to 10:00 / 00:00 to 08:00 | 03:00 to 11:00 / 01:00 to 09:00 | LatAm operators come online near handoff: morning deposit and withdrawal questions. |
| London | 06:00 to 14:00 | 06:00 to 14:00 | European open. Fiat rails and banking partners start moving. |
| Dubai | 09:00 to 17:00 | 10:00 to 18:00 | Full business day inside the window. |
| Bangkok, Jakarta | 12:00 to 20:00 | 13:00 to 21:00 | Full business day and evening retail peak. |
| Singapore, Hong Kong, Manila | 13:00 to 21:00 | 14:00 to 22:00 | Full business day and evening peak. Operator staff escalate in real time. |
| Tokyo, Seoul | 14:00 to 22:00 | 15:00 to 23:00 | Afternoon and evening retail peak. |
EDT runs from March to November. The shift is pinned to New York time, so regions without matching daylight saving shift by one hour at each US change.
Most white-label exchange vendors sell a feature list. Regulated buyers also score uptime and incident handling, which is where TechOps shows up in a sales cycle.
| Vendor | Positioning | Gap vs AlphaPoint | Operations angle |
|---|---|---|---|
| AlphaPoint | Institutional white-label exchange, brokerage, perps, treasury. SOC 1 / SOC 2. Bank-affiliated customer (Wenia). | Benchmark. | Managed hosting across mixed infrastructure; the run quality is part of the product. |
| Shift Markets | Modular stack and operator back office, faster launches. | No comparable public audit or bank-affiliated customer found. | Competes on speed to launch. |
| B2Broker (B2Trader) | Liquidity, CFD / forex and crypto package for brokers. | Broker-first buyer; different compliance profile. | Liquidity bundled with the platform. |
| ChainUP | Asia-based, low price, many modules. | Less US and regulated-bank credibility. | Price pressure in Asian deals. |
| HollaEx | Open-source exchange kit for small operators. | Self-hosted by the operator; outside the institutional segment. | Operators carry their own uptime. |
| Modulus, OpenDAX (Yellow) | Older US vendor; open-source stack now pivoted to the Yellow network. | Smaller footprint or less vendor support (my assessment). | Migration candidates when a platform stalls. |
| Nasdaq Market Technology | Exchange-grade matching and surveillance for large venues. | Priced for national exchanges (my assessment). | Sets the uptime expectation that institutional buyers bring. |
Four public gateways answer the APEX frame format. Since 14 September 2026 the demo probes each one every 60 seconds from Cloudflare (Ping, GetInstruments, SubscribeLevel1) and records the results. Public, unauthenticated calls only. First readings:
| Gateway | Instruments | Not Running | Connect / Ping | Read |
|---|---|---|---|---|
| apexapi.bitazza.com | 307 | 34 (Stopped) | 40 ms / 4 ms | Delisted pairs stay Stopped. State-based alerting pages forever; baseline diff pages only on change. |
| api.coinext.com.br | 88 | 44 | 756 ms / 250 ms | Named AlphaPoint customer. Half the instrument list is not Running and BTC/BRL showed a 327 bps spread with the last trade 95 minutes old: a thin book is normal here, so staleness thresholds must be per venue. |
| api.foxbit.com.br | 134 | 0 | 740 ms / 242 ms | Latency tracks distance from the probe. A baseline per gateway and vantage is needed before any latency threshold means something. |
| api.ndax.io | 90 | 0 | 792 ms / 256 ms | An earlier run from the same edge measured 1,678 ms connect and 1,099 ms ping. One sample is noise; alert on sustained deviation. |
Record normal per gateway (instrument states, RTT band), then alert on departure from it.
Browser and edge disagree: network or WAF. Both fail: gateway. The console shows the verdict on each card.
An old Level 1 timestamp on a thin THB pair is normal. The same age on BTC/USD is an incident.
| JD duty | How I would do it | Detailed in |
|---|---|---|
| System uptime and integrity of customer systems | Synthetic gateway probes per customer from two vantages, baseline diffs, P1 to P3 definitions tied to customer impact. | ★, §05 |
| Documenting designs, training materials | One runbook per alert: symptom, first safe command, escalation owner, rollback. New hires shadow with the runbook open. | §06, §08 |
| Monitoring hardware, software, OS | Read-only host snapshot (disk, inodes, memory, failed units, NTP, OOM, ZFS, storage, certs) collected on a schedule and diffed day to day. | §07 |
| Linux, Proxmox and Docker installs, upgrades, maintenance | Patches in change windows through config management; pending reboots and security updates reported at every handoff. | §05, §07 |
| Automation with dev-ops | Turn each repeated night check into an Ansible role or scheduled job, then delete the manual step from the runbook. | §07 |
| Protection of data assets, security policy | Flag exposed MySQL, Redis, Docker API and Proxmox UI ports; key-only SSH; least-privilege tooling that reads and never writes. | §06 |
| Testing, troubleshooting, modifying systems | Measure first, contain, then fix. Every change logged with NY and UTC timestamps for audit evidence. | §06 |
| Working with TechOps, IT, IT Security, sysadmins | Written handoff at 09:00 EST with open items by severity, actions taken and the next owner. | §08 |
A proposal to align with the existing SLAs on day one. Severity follows customer impact first and host metrics second.
| Sev | Definition | Examples | Response |
|---|---|---|---|
| P1 | Trading, deposits or withdrawals impaired for a customer, or data integrity at risk. | Gateway down from two vantages; pair leaves Running unplanned; crossed book; replica SQL thread stopped; ZFS pool degraded; cert expiring within 7 days. | Page on-call, open incident, customer update within the agreed window. |
| P2 | Degraded or one failure away from P1. | Storage above 85%; replication lag rising; OOM kills; NTP unsynchronised; WAN failover active; state table above 80%. | Ticket, fix or mitigate on shift, flag in handoff. |
| P3 | Hygiene with no current impact. | Pending security updates; reboot required; swap creeping; noisy journal. | Change ticket for the next window. |
The overnight engineer is alone. The judgment is knowing which fixes are safe to make without an owner awake.
| Situation | Safe on my own | Needs the owner |
|---|---|---|
| MySQL replica SQL thread stopped | Capture Last_SQL_Error and relay position, tell customer-facing staff that replica reads are stale. | Skipping the event or promoting the replica. A skipped row on a trading ledger is a silent balance mismatch. |
| Proxmox cluster lost quorum | Check corosync links and node reachability, stop further changes. | pvecm expected 1 or any forced quorum: split-brain risk. |
| pfSense CARP roles split | Confirm which node passes traffic, check pfsync and peer health. | Toggling persistent maintenance mode or rebooting the master. |
| Disk full on a host | Measure with du and docker system df; rotate logs; clear build cache. | Deleting volumes, VM disks or backups. Named volumes can look unused while holding data. |
| Container crash loop | Read logs, exit code, OOM flag and cgroup memory.events; roll back to the previous image if the runbook allows. | Raising memory limits on a shared node, which moves the pressure to every other guest. |
| Instrument halted | Verify from two vantages, check the maintenance calendar, notify. | Resuming trading on the instrument. |
| Wallet signer VM stopped | Confirm state and alert. | Starting it. Signing infrastructure follows the custody procedure. |
Ranked by pages saved per hour of work. The first two exist in the demo as working prototypes.
| # | Item | What it does | Effort | Status |
|---|---|---|---|---|
| 1 | Host snapshot role | ap-health.sh as an Ansible role on every host; outputs stored per day so triage becomes a diff. | S | Script in demo |
| 2 | Gateway probe with baseline | Scheduled Ping / GetInstruments / GetLevel1 per customer; alerts on change from baseline. | S | Running in demo |
| 3 | Certificate expiry sweep | Every customer hostname checked daily; ticket at 30 days, page at 7. | S | Planned |
| 4 | Replica health probe | IO / SQL thread state and lag trend exported as metrics, with an absent-data alert. | M | Planned |
| 5 | pfSense state and gateway export | State table usage, dpinger loss and CARP role scraped into monitoring. | M | Script in demo |
| 6 | Handoff template from the incident log | Open items by severity, timeline in NY and UTC, generated at 09:00. | S | Prototype in demo |
| 7 | Patch-window report | Pending security updates and reboot flags per host, ready for the change ticket. | S | Planned |
Built from the public job posting (2026), AlphaPoint public pages, press coverage and live probes of public APEX gateways on 2026-09-14. Probes used unauthenticated read calls only. Hosting inferences are mine and labelled as such. The Overnight Console is my own tool built for this application; it has no connection to AlphaPoint or customer systems. Unsolicited interview homework; happy to walk through any section.
Leadership: alphapoint.com/leadership
API docs: AlphaPoint University · apex-api
SOC 1 / SOC 2: AlphaPoint blog
Hosting history: Microsoft case study
Funding: The Block, 2020
Wenia: AlphaPoint press · Green-X: case study
NDAX: Yahoo Finance · Coinext: case study
Perps: PR Newswire, Aug 2025
POLYX: PR Newswire, Oct 2025
Vendor landscape: Finance Magnates, 2026
Independent TechOps homework for the AlphaPoint Tech Ops Engineer (overnight) role · 2026 · edwardtay.com