Skip to content
Omar Sharkeyeh

CVSelected workNun Pirat

Nun Pirat: the platform under an Arabic learning app

Nun Pirat teaches children, families and school classes to read and write Arabic. I build the app, and I designed and built the platform it runs on: the infrastructure, the delivery pipeline and the monitoring. This page is about the platform.

My role
Design, build and run the platform
Time
App since 2024, platform 2026
Environments
Staging and production
Code
15 platform repositories, one pipeline
The app at nunpirat.com. Everything below is what keeps this page fast, up and safe to change.

What it had to do

Families use the app in the evening and schools in the morning, so it has to be quick at both peaks and should not go down in between. The number of learners was going to grow, and the platform had to grow with it without a rebuild. There is no operations team, so the routine work had to run by itself. It should not be tied to one cloud. And shipping a change had to be safe enough to do on an ordinary Tuesday afternoon.

The stack

The platform is built in four layers. Each one only relies on the layer below it: the app does not care which servers it runs on, and only the bottom layer knows which cloud is underneath.

Application

What the learners use

LaravelReactPostgreSQLRedisTypesense searchbackground workers

Platform services

What every app on it gets

certificatespublic and private DNStraffic routingdatabase operatorshared storageVPNsafe reboots

Kubernetes

Where everything runs

staging clusterproduction clusterautoscalingself-updating

Foundation

The only cloud-specific part

networksprojectsaccess for the pipelines
The top three layers are plain Kubernetes, Helm and Terraform. Moving to another provider means rewriting the foundation; the platform and the app stay as they are.

Everything runs from code

No server on this platform was set up by hand. Clusters, networks, DNS records, certificates, databases and access rights are written down in Terraform and applied by the pipeline. A new environment is the same code run again, which is why staging and production only differ in size. A change to the infrastructure goes through review like any change to the app, with a preview of exactly what will change. And because the code is the description of the platform, it does not drift out of date the way a wiki page does.

Automated from commit to production

All repositories, app and platform, share one set of pipeline templates, so a change takes the same five steps wherever it is made. One of them needs a person. Tests run at three of them; the testing section below shows which.

  1. Merge requestA developer proposes a change.
  2. Build and testautomaticImage built, unit and feature tests, code and infrastructure checks.
  3. StagingautomaticDeployed straight away, then the end-to-end tests run against it.
  4. Approvalone clickOnly offered once every test has passed.
  5. ProductionThe same release goes live and is checked once more.
Safety netDatabase changes run before the new version starts. If one fails, the release stops and users keep the version that was running.
Same release everywhereStaging and production get the same image and the same infrastructure code, so what was tested is what goes live.
Infrastructure changes take the same path: a change to a cluster or a DNS record is previewed, applied to staging and approved for production like an app release.

What else runs on its own

Updates for the components the platform depends on arrive as merge requests, so they get the same review and the same tests as any other change. The operating system layer under the app is rebuilt every night, which means security patches reach production with the next release. And the test environment keeps office hours:

It starts at nine, hibernates at six, and costs next to nothing overnight.

Automated testing

The tests are part of the pipeline. A release has to get through three layers of them before anyone is asked to approve it, and production is checked once more after it goes live.

  1. End-to-end tests

    Nine scenarios in a real browser: registering as an adult or a parent, signing in with a one-time code, child accounts, and checkout with Stripe and PayPal in sandbox mode, including the e-mails that go out.

    on every staging release
  2. Feature tests

    Pages and interfaces against a real database: sign-in, billing, lessons, notifications, privacy, search engine tags.

    on every merge request
  3. Unit tests and static checks

    Business rules, frontend components, types and code style; the infrastructure code is validated and previewed. Cheap and fast, so they run on every commit.

    on every commit
backend tests
about 1,300
frontend test files
40
end-to-end scenarios
9
After go-liveA check confirms that every part of the app runs the new version and that the pages render. If not, the release is marked as failed.
Under loadLoad tests simulate many users at once and show how much traffic the platform can take before the real traffic arrives.
The many quick tests at the bottom catch most mistakes within minutes of a commit. The few slow ones at the top show that the whole thing works together.

360° monitoring

An uptime check can only tell you whether the site answers. On this platform every part reports what it is doing, so you can also see how well it is doing, what happened, where a request went and where the time was spent.

Sources

  • Web app
  • Background workers
  • Scheduled tasks
  • Live updates
  • Cluster and nodes
MetricsHow is it doing?Prometheus
LogsWhat happened?Loki
TracesWhere did the request go?Tempo
ProfilesWhere is the time spent?Pyroscope
OutsideIs the site up for users?Uptime Kuma

Grafana

One dashboard
  • a map of every service
  • signals linked to each other
  • the release on every signal
  • alerts
Everything reports through one open standard, OpenTelemetry. The signals are stored by kind and meet again in one dashboard, which opens on a map of all services coloured by health.

From alert to cause

Because the signals are linked and each one carries the release it came from, finding out why something got slow is a short walk instead of a search through several tools:

  1. AlertResponse times are up.
  2. TraceOne slow request, step by step.
  3. LogsWhat happened during that request.
  4. ReleaseThe version that introduced it.

When something goes wrong

Why it is built this way

So that a small team can run it. Nobody has to remember a manual step, a failed release does not reach users, and when something does go wrong the data to explain it is already there.

Back to the CV