CVSelected workNun Pirat
Nun Pirat: the platform under an Arabic learning app
Nun Pirat teaches children, families and school classes to read and write Arabic. I build the app, and I designed and built the platform it runs on: the infrastructure, the delivery pipeline and the monitoring. This page is about the platform.
- My role
- Design, build and run the platform
- Time
- App since 2024, platform 2026
- Environments
- Staging and production
- Code
- 15 platform repositories, one pipeline
What it had to do
Families use the app in the evening and schools in the morning, so it has to be quick at both peaks and should not go down in between. The number of learners was going to grow, and the platform had to grow with it without a rebuild. There is no operations team, so the routine work had to run by itself. It should not be tied to one cloud. And shipping a change had to be safe enough to do on an ordinary Tuesday afternoon.
The stack
The platform is built in four layers. Each one only relies on the layer below it: the app does not care which servers it runs on, and only the bottom layer knows which cloud is underneath.
Application
What the learners use
Platform services
What every app on it gets
Kubernetes
Where everything runs
Foundation
The only cloud-specific part
Everything runs from code
No server on this platform was set up by hand. Clusters, networks, DNS records, certificates, databases and access rights are written down in Terraform and applied by the pipeline. A new environment is the same code run again, which is why staging and production only differ in size. A change to the infrastructure goes through review like any change to the app, with a preview of exactly what will change. And because the code is the description of the platform, it does not drift out of date the way a wiki page does.
Automated from commit to production
All repositories, app and platform, share one set of pipeline templates, so a change takes the same five steps wherever it is made. One of them needs a person. Tests run at three of them; the testing section below shows which.
- Merge requestA developer proposes a change.
- Build and testautomaticImage built, unit and feature tests, code and infrastructure checks.
- StagingautomaticDeployed straight away, then the end-to-end tests run against it.
- Approvalone clickOnly offered once every test has passed.
- ProductionThe same release goes live and is checked once more.
What else runs on its own
Updates for the components the platform depends on arrive as merge requests, so they get the same review and the same tests as any other change. The operating system layer under the app is rebuilt every night, which means security patches reach production with the next release. And the test environment keeps office hours:
Automated testing
The tests are part of the pipeline. A release has to get through three layers of them before anyone is asked to approve it, and production is checked once more after it goes live.
- End-to-end tests
Nine scenarios in a real browser: registering as an adult or a parent, signing in with a one-time code, child accounts, and checkout with Stripe and PayPal in sandbox mode, including the e-mails that go out.
on every staging release - Feature tests
Pages and interfaces against a real database: sign-in, billing, lessons, notifications, privacy, search engine tags.
on every merge request - Unit tests and static checks
Business rules, frontend components, types and code style; the infrastructure code is validated and previewed. Cheap and fast, so they run on every commit.
on every commit
- backend tests
- about 1,300
- frontend test files
- 40
- end-to-end scenarios
- 9
360° monitoring
An uptime check can only tell you whether the site answers. On this platform every part reports what it is doing, so you can also see how well it is doing, what happened, where a request went and where the time was spent.
Sources
- Web app
- Background workers
- Scheduled tasks
- Live updates
- Cluster and nodes
Grafana
One dashboard- a map of every service
- signals linked to each other
- the release on every signal
- alerts
From alert to cause
Because the signals are linked and each one carries the release it came from, finding out why something got slow is a short walk instead of a search through several tools:
- AlertResponse times are up.
- TraceOne slow request, step by step.
- LogsWhat happened during that request.
- ReleaseThe version that introduced it.
When something goes wrong
- A change breaks registration. The end-to-end tests on staging fail, the release is never offered for approval, and production carries on as before.
- A release contains a database change that fails. It fails before the new version starts, the release stops, and users stay on the previous version while the pipeline shows what went wrong.
- Pages get slower after a release. The dashboard shows the rise, the traces show which step slowed down, and the release on every signal points to the change behind it.
Why it is built this way
So that a small team can run it. Nobody has to remember a manual step, a failed release does not reach users, and when something does go wrong the data to explain it is already there.