In short for the platform team

What it isColloq is a collaborative Jupyter classroom: a web app, a runtime broker and a separate Pod with a Python kernel for each room. It goes into your cluster as a Helm chart, oci://ghcr.io/colloq-edu/charts/colloq, in one namespace. Open source, MIT license.
What runsThe app Deployment, one replica: the database is SQLite. The broker Deployment, one replica: the only thing that talks to the Kubernetes API. A Pod and a Service for each running room. A second Pod and Service per room for students' personal notebooks. Short-lived competition Pods: submissions, the metric, package preparation. Room and competition Pods are created by the broker, not by the chart.
What it needsOne namespace with namespace-admin rights; the broker's Role from Permissions; a StorageClass — ReadWriteOnce required, ReadWriteMany preferred; an ingress controller and a certificate; a registry mirror with the release's images.
What it never needscluster-admin or cluster-scoped objects: Namespace, PersistentVolume, ClusterRole, CRD, DaemonSet, RuntimeClass. Privileged containers, hostPath, hostNetwork or hostPID. Root: every container runs as uid 1000. Internet access from the cluster.
NetworkInbound: only the ingress controller to the app, port 3000. Inside the namespace: app → broker, app and broker → rooms. Outbound: broker → Kubernetes API, Pods → cluster DNS; optionally app → the AI model and an outbound proxy. Students' code gets DNS only; competition submissions get nothing.
DataTwo PVCs: colloq-data — the SQLite database, keys, class page files, competition data and submissions, database snapshots; colloq-workspace — room files.
BackupsVolume snapshots with your own tools (Velero, CSI), plus consistent database snapshots the app writes into the colloq-data volume itself once a day.
Students' and teachers' browsersThe cluster's ingress controller, TLSApp · Deployment colloq-app, 1 replicaBroker · Deployment colloq-runtime, 1 replica
Room A · Pod and ServiceRoom B · Pod and ServiceThe room's personal notebooksCompetition submission

The app — a Node.js server with the built web interface and an SQLite database — keeps everything on two volumes and never calls the Kubernetes API. When someone joins a room or runs a cell, the app asks the broker for the room's Pod. The broker creates the Pod and a ClusterIP Service from a template fixed in its code, waits for Jupyter and hands the app its address and token; from then on the app talks to the room's Jupyter directly, on port 8888. A room's Pod lives while anyone is in the room and for 2 hours after everyone has left. Room Pods survive restarts and upgrades of the app and the broker: the new broker finds them by their labels, and Python variables stay.

What we assume about your cluster

This is a checklist for your team: go through it before installing. The answer to each item is either "yes, that's us" or a chart value to set.

WhatWe assumeIf yours differs
Kubernetes version1.30 or newer, upstream-compatible: vanilla, Deckhouse and the like. Changing a running room's memory and cores without a restart needs 1.33 or newer and the pods/resize permission.Older than 1.33, or no such permission: runtime.inPlaceResize: false. The rule is dropped from the Role, resizing a live room answers "unsupported" rather than failing, and the new number goes to the room's next Pod.
NamespaceWe get one namespace with namespace-admin rights, and it holds one Colloq installation: object names are fixed (colloq-app, colloq-runtime, …). The chart creates no cluster-scoped objects.Your team creates the namespace, its labels and its quota. The chart creates the Role and RoleBinding itself, and Kubernetes won't let anyone grant a verb they don't hold: whoever installs the chart needs every verb in Permissions.
Pod Security and policiesPod Security Admission restricted is enforced, and Kyverno or Gatekeeper require requests and limits on every container, probes, runAsNonRoot and readOnlyRootFilesystem, forbid privileged, hostPath and hostNetwork, admit images only from the internal registry and only by digest, and require the standard app.kubernetes.io/* labels.Every Colloq Pod passes restricted; details and exceptions are under Security. commonLabels adds your labels to every chart object and every broker Pod; rooms.podLabels and rooms.podAnnotations add them to broker Pods only.
Network policiesNetworkPolicy is enforced (Calico, Cilium), and the namespace may be default-deny.The chart declares every flow it needs, ingress and egress; they are listed under Network. Say where the ingress controller lives (networkPolicy.ingressController.from), where cluster DNS lives (networkPolicy.dns.to), and the API servers' addresses (networkPolicy.kubeApiServer).
Ingress and TLSIngress through ingress-nginx, with a configurable class and annotations. TLS: the organisation's certificate from an existing Secret or from cert-manager. Users come from the corporate network, over VPN or, when students are external, from the internet.Another controller: ingress.className and ingress.annotations with the same meaning — request bodies of 256 MB or more, timeouts of an hour or more, no response buffering. Or ingress.enabled: false: publish the Service colloq-app:3000 with your own object and set config.publicUrl.
Sign-inStudents join with the class link, teachers sign in with a personal link. A proxy with SSO that forwards a signed JWT — Teleport application access, oauth2-proxy or Pomerium in front of your IdP — may stand in front of Colloq.config.sso.*: teachers sign in with their corporate account, students who have one join under their own name; see Sign-in through your SSO. Without such a proxy, nothing to set.
StorageA StorageClass with ReadWriteOnce always exists; one with ReadWriteMany (CephFS) may.By default the chart assumes ReadWriteOnce: everything that mounts Colloq's volumes runs on one node. With ReadWriteMany, set persistence.accessMode: ReadWriteMany and room Pods spread over the cluster; see Storage.
No internetThe cluster has no internet access. Images come through a registry mirror (Harbor, Nexus), PyPI packages through a Nexus mirror, and the Oracle through an internal model with an OpenAI-compatible API, or it stays off. There may be an outbound HTTP proxy and an internal CA.global.imageRegistry — the mirror for every image, kernel images included; dependencies.indexUrl and dependencies.mirrorEgress — the PyPI mirror; config.ai.baseUrl — the model; outbound.httpsProxy, outbound.noProxy, outbound.extraCa. See Images and the registry mirror.
Secrets and GitOpsSecrets come from Vault through External Secrets, and deployment goes through ArgoCD.Every secret can come from an existing Secret: config.existingSecret, secrets.runtimeToken.existingSecret, secrets.roomSecret.existingSecret, metrics.existingSecret. With those set, the chart renders no random value, so ArgoCD sees no perpetual drift.
Monitoring and logsPrometheus Operator; logs are collected from stdout and stderr.metrics.enabled turns on /metrics with a token, metrics.serviceMonitor.enabled adds a ServiceMonitor; without the Prometheus Operator, scrape /metrics your own way. config.logFormat: json writes one JSON object per line.
BackupsVolume snapshots: Velero, CSI.The app itself writes consistent database snapshots into the colloq-data volume (config.dbSnapshotHours, config.dbSnapshotKeep), so a volume snapshot always carries a whole copy of the database.
GPUGPU nodes are optional.GPU rooms take a RuntimeClass, a nodeSelector and tolerations: gpu.runtimeClassName, gpu.nodeSelector, gpu.tolerations; see Operations.

Permissions

Only the broker has rights in the cluster: the Role and RoleBinding colloq-runtime in this namespace. No ClusterRole, no access to nodes, Secrets, Deployments or other namespaces.

ResourceVerbsWhy
podsget, list, create, deleteRoom, personal-notebook and competition Pods. list is for the census: a restarted broker finds the running Pods by their labels and adopts them instead of creating them again.
servicesget, list, create, deleteA ClusterIP Service for each room and its personal notebooks — the stable name the app reaches Jupyter by — and for the package-preparation proxy. A room deleted for good leaves a tombstone: a headless Service with no selector and no ClusterIP, so that its ID can't be opened again.
pods/resizepatchChange a running room's memory and cores without a restart (Kubernetes 1.33+): the patch touches only memory and cpu of the kernel container. Optional: with runtime.inPlaceResize: false the rule is gone.
persistentvolumeclaimsgetBefore a competition, check that the data volume is in place.
networkpolicies (networking.k8s.io)getBefore a competition, check that the submissions' isolation policies exist. Without them submissions are refused, and the teacher sees why.

There is no patch or update on Pods and Services: the broker doesn't edit live objects — it changes only a room's resources through pods/resize, and deletes and recreates everything else, with a UID precondition. RBAC can't constrain the fields of a Pod, so the Pod templates are fixed in the broker's code, and its HTTP API takes only a room ID, an environment from the catalog, an image revision, memory and cores: no arbitrary template, no hostPath, no image outside the catalog. Restricted Pod Security and your policies hold the rest.

ServiceAccountAPI tokenWho
colloq-appnoThe app. It never calls the Kubernetes API.
colloq-runtimeyes: the broker's Pod mounts a projected token that the kubelet rotatesThe broker; bound to the Role above.
colloq-kernelnoRoom and personal-notebook Pods: students' code holds no API credentials.
the namespace defaultnoCompetition Pods: no token is mounted.

All three ServiceAccounts set automountServiceAccountToken: false; only the broker's Pod turns the token on.

Security

Pod Security. Every Colloq Pod — the app, the broker, rooms, personal notebooks, competitions — passes the restricted profile: runAsNonRoot, uid and gid 1000, seccomp RuntimeDefault, allowPrivilegeEscalation: false, capabilities.drop: [ALL], readOnlyRootFilesystem: true. Only volumes are writable: emptyDir for /tmp, the home directory and /dev/shm (in memory: 64 MiB, 1 GiB for GPU rooms), and the PVCs. Your team puts the pod-security.kubernetes.io/enforce: restricted label on the namespace: the chart doesn't own it.

Policy engines.

  • Requests and limits are set on every container, including the init container and the sidecar of competition Pods. A room kernel's requests equal its limits (QoS Guaranteed): memory and cores are reserved in full, plus 2 GiB of ephemeral storage.
  • Probes. The app: HTTP startupProbe and livenessProbe on /api/livez, readinessProbe on /api/readyz. The broker: TCP 8787. Room and personal-notebook kernels: TCP readiness and liveness probes on 8888; the liveness probe acts only after five minutes of refused connections, since a Pod with restartPolicy: Never ends with the class's variables when it does. Submission containers run to completion and have no probes; their exporter and the package proxy have readiness probes. If a probe rule doesn't fit run-to-completion Pods, exempt them by the label app.kubernetes.io/managed-by: colloq-runtime.
  • Labels. Chart objects carry app.kubernetes.io/name: colloq, instance, component, version, part-of, managed-by: Helm and helm.sh/chart. Pods the broker creates carry app.kubernetes.io/name, instance and part-of, app.kubernetes.io/managed-by: colloq-runtime and colloq.dev/role: kernel (the personal-notebook Pod also has colloq.dev/kernel: own), competition-job, competition-resolver, competition-proxy. The app.kubernetes.io/* and colloq.* keys belong to the chart and the broker and can't be overridden.
  • Digests. The chart always renders the app and broker images as registry/path@sha256:… — never a tag, never latest. Kernel images in the catalog are accepted only in the same form: neither the chart nor the broker lets a tag-only reference through.
  • Registry. global.imageRegistry rewrites the registry of every image, including the catalog's kernels, the competition exporter and the package proxy.
  • Mutating webhooks. After creating a room's Pod the broker checks it against what it asked for. If a webhook changed its containers, volumes or security settings — injected a sidecar, rewrote imagePullPolicy — the broker refuses the Pod and the room doesn't start; the broker's log names the fields that changed. Exempt Pods labelled app.kubernetes.io/managed-by: colloq-runtime from such webhooks; for Istio, rooms.podLabels: {sidecar.istio.io/inject: "false"} is enough. Placement the cluster fills in by itself — PodNodeSelector, PodTolerationRestriction, priority, RuntimeClass, Kueue — doesn't disturb the check.

What code runs as whom.

PodImage and codeAPI tokenVolumesNetwork
Appcolloq-app: the Colloq server on Node.jsnocolloq-data and colloq-workspacein from the ingress controller; out to the broker, rooms, DNS and, optionally, the AI model and a proxy
Brokercolloq-runtime: the broker on Node.jsyesnonein from the app; out to the Kubernetes API, rooms, submissions and DNS
Roomthe room environment's colloq-kernel: Jupyter and students' codenoits own folder of colloq-workspacein from the app and the broker on 8888; out: DNS only
Personal notebookscolloq-kernel: students' code in personal notebooks, no GPUnoits room's folder, at the same pathas a room
Submissioncolloq-kernel: the entrant's notebook or the metric; colloq-runtime: the result exporternofolders of colloq-data, inputs read-onlynone; in: only the broker to the exporter on 8765
Package preparationcolloq-kernel: pip downloads entrants' own packagesnofolders of colloq-dataonly to its own proxy on 3128, and DNS
Package proxycolloq-runtime: a CONNECT proxynononein from preparation; out: DNS and the package index — public PyPI on 443 or your mirror

Isolation between rooms. Each room has a Pod of its own: its own processes, its own emptyDir, its own Jupyter token — the broker derives it from the room secret with an HMAC and gives it to the app only. The Pod mounts only its own folder of the colloq-workspace volume (subPath); other rooms' files aren't in it. The network policy lets only the app and the broker reach a room's port 8888 and lets nothing out of a room but DNS: neighbouring rooms, the app and the API are out of reach of students' code. Inside a room, participants share the kernel, the files and the Linux user — by design: it is a shared workspace. Standard NetworkPolicy doesn't guarantee that a Pod can't reach services on its own node: test it from a room (see the checks after installing) and close it with your CNI's host policies if needed.

Secrets.

SecretWhat it holdsWho reads it
config.existingSecret or colloq-app-configKeys are app settings: OPENAI_API_KEY, SESSION_SECRET, a proxy with a password in HTTPS_PROXY, or any other settingthe app
colloq-runtime-auth, key runtime-tokenThe Bearer token the app uses with the brokerthe app and the broker
colloq-room-secret, key room-secretThe key the broker derives rooms' Jupyter tokens and submission export tokens fromthe broker only
colloq-metrics, key metrics-tokenThe Bearer token for /metrics, when metrics are onthe app and Prometheus
global.imagePullSecretsThe mirror's .dockerconfigjson; broker Pods get the first one toothe kubelet

Each can come from an existing Secret — for example one External Secrets creates from Vault. When the broker token, the room secret or the metrics token are not given, the chart generates them: under the helm CLI once, reading them back from the cluster with lookup afterwards. ArgoCD and other tools that render the chart with helm template can't read them back and would generate new values on every sync, restarting rooms — set existingSecret there. The broker token and the room secret are 43 to 128 characters of [A-Za-z0-9_-] from at least 32 random bytes: openssl rand -hex 32. Keep the secrets stable: a new SESSION_SECRET invalidates every link and sign-in handed out, and a new room secret replaces room Pods at their next start, losing Python variables. Without SESSION_SECRET, the app generates one once and keeps it in the colloq-data volume. When you move from a generated secret to existingSecret, copy its value into the new Secret first: the chart deletes its own.

Secrets live in the colloq-data volume too: the setup token, the teachers' sign-in keys in the database, the model key if it was entered in the panel. Keep snapshots of that volume as private as Secrets.

Network

Every flow Colloq needs. The chart declares them as network policies, ingress and egress, so the installation works in a default-deny namespace too.

FromToPortWhy
Ingress controllerAppTCP 3000All user traffic: pages, API, WebSocket, server-sent events. networkPolicy.ingressController.from says who is admitted (by default the ingress-nginx Pods in the ingress-nginx namespace). A controller on the host network arrives from node addresses: give an ipBlock of your nodes then.
AppBrokerTCP 8787Start, resize and stop a room's Pod, start a submission. Requests carry a Bearer token.
AppRoom and personal-notebook PodsTCP 8888Jupyter: running cells, the terminal. With the room's token.
BrokerRoom and personal-notebook PodsTCP 8888Check that Jupyter is up and accepts the token.
BrokerSubmission and package-preparation PodsTCP 8765Collect the result from the exporter.
BrokerKubernetes API server443 and 6443Create and delete Pods and Services. The network policy sees the address after DNAT, that is the API servers' own addresses, so by default any address on these ports is allowed. Narrow networkPolicy.kubeApiServer.cidrs to the output of kubectl get endpointslices -n default -l kubernetes.io/service-name=kubernetes. On Cilium the API server has an identity of its own that an ipBlock doesn't match: turn on networkPolicy.kubeApiServer.ciliumEntity, and the chart adds a CiliumNetworkPolicy in this namespace.
Every Colloq Pod except submissionsCluster DNSUDP and TCP 53Service names. networkPolicy.dns.to says where DNS is (by default k8s-app: kube-dns in kube-system). Rooms can lose DNS too: rooms.network: none.
Package preparationIts own package proxyTCP 3128pip goes only through a proxy that admits one package index.
Package proxyPyPI mirror or public PyPI443 or the mirror's portIndex and files of competition entrants' own packages. A policy knows addresses, not names: the mirror's addresses and ports go into dependencies.mirrorEgress. In a closed network turn off the general rule for public addresses on 443: dependencies.publicEgress: false.
AppAI model, outbound proxy, GitHubthe model's or proxy's portOptional: the Oracle and notebook import from GitHub. By default the app may go nowhere outside: addresses and ports go into networkPolicy.app.egress.
PrometheusAppTCP 3000, /metricsOptional, with metrics.enabled: networkPolicy.monitoring.from says who is admitted.
Nodes (kubelet)Registry mirrorTCP 443Images. Not described by network policies.

If the cluster runs NodeLocal DNSCache, DNS queries go to a node-local address rather than to the kube-dns Pods: add that address to networkPolicy.dns.to, for example ipBlock: {cidr: 169.254.20.10/32}.

Students' code doesn't get out: room and personal-notebook Pods have no egress but DNS, and rooms.network: none takes that away too. So pip install in a cell doesn't work: libraries come in kernel images; see Images and the registry mirror. Competition submissions get no network at all, not even DNS, and that is checked before any code runs: an init container tries to get out and lets the submission run only after three refusals in a row; if the network answers, the submission fails.

The client address: TRUSTED_PROXIES=private. Per-address caps — 60 newcomers to a room per 10 minutes, 20 Oracle questions a minute, competition sign-ins and joins — must see the student's address, not the ingress controller's. The app believes X-Forwarded-For only from connections whose address is in TRUSTED_PROXIES; the chart sets private (config.inbound.trustedProxies) — every private range, wherever the controller's Pod comes from. That is safe because only the ingress controller is admitted to the app's port 3000: the network policy lets no one else in, students' code included, so nobody inside the cluster can forge the header. The client is the rightmost address in the header that is not itself trusted; when every hop is trusted — and on a corporate network users' addresses are private too — the leftmost stands, which is the one the controller wrote. So the controller must replace the header rather than append its own address to one the client sent: ingress-nginx does that by default, while use-forwarded-headers and compute-full-forwarded-for are off. If a load balancer in front of the controller rewrites the source address, every student looks like one address: preserve the client address (externalTrafficPolicy: Local on the controller's Service, or PROXY protocol). If many people come from one address — through a NAT or a VPN concentrator — list it in config.inbound.sharedAddresses: per-address caps don't apply to it, while per-room and per-person caps still do.

ingress-nginx. The chart puts these annotations on its Ingress (ingress.annotations; a key set to null drops an annotation):

nginx.ingress.kubernetes.io/proxy-body-size: 256m       # competition data: up to 200 MB in one request
nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"  # WebSocket and server-sent events live for hours
nginx.ingress.kubernetes.io/proxy-send-timeout: "3600"
nginx.ingress.kubernetes.io/proxy-buffering: "off"      # events must arrive one by one

Without the first line ingress-nginx accepts no more than 1 MB, and uploading competition data or a file to a room gets a 413. The controller passes WebSocket (/collab/, /control/, /file/) by itself; the server pings every 25 s, and server-sent event streams send a heartbeat every 20 s and the header X-Accel-Buffering: no. The app serves the root of its own hostname and won't work from a sub-path: the chart's Ingress is by the host ingress.host, path /. A file uploaded to a room is up to 50 MB by default (config.maxUploadMb); raise that and raise proxy-body-size too.

TLS and the address. The organisation's certificate is an existing Secret of type kubernetes.io/tls: ingress.tls.secretName, colloq-tls by default. Or cert-manager: ingress.tls.certManager.clusterIssuer adds the annotation with your issuer. PUBLIC_URL is the external address every link Colloq hands out is built from, the owner link included; by default the chart takes https://<ingress.host>, and config.publicUrl sets another. When the controller says X-Forwarded-Proto: https, cookies get the Secure flag and the app sends Strict-Transport-Security. ingress-nginx adds that header itself by default; to avoid sending it twice, turn the app's off: config.inbound.hsts: false.

The ingress controller's access logs contain room addresses, staff sign-in keys (/admin/k/…, /admin/t/…) and room tokens in WebSocket addresses (?token=…). Keep them as private as the database.

Sign-in through your SSO

If a proxy in front of Colloq checks the person itself and forwards a signed JWT with every request — Teleport application access, oauth2-proxy or Pomerium in front of your IdP — Colloq accepts that sign-in. The signature is checked against the proxy's public keys (JWKS), an expiry and the audience are required, and the issuer is checked when set. The audience is required because one proxy signs the tokens of all its applications: without it, a token issued for another application would sign in here too; * accepts any. Only asymmetric algorithms are accepted (RS, PS, ES, EdDSA); none and HMAC are always refused. A request without a token, or with one that does not verify, is served as before: links keep working, and the proxy cannot lock a class out.

  • Teachers. A token whose email is on the staff list signs in as that teacher: Colloq issues the ordinary cookie, and the panel, the room and the sockets see the signed-in teacher. The owner still adds colleagues by email, but no longer needs to send them a link. The proxy roles in config.sso.teacherRoles and config.sso.ownerRoles add a person to the list on their first visit; they never change the role of someone already on it. With these roles the proxy decides the list: someone removed in the panel comes back while the proxy still gives the role, so take the role away at the proxy. Only an owner role claims an unclaimed instance. A cookie a shared browser kept from the previous person never outranks the person the proxy vouches for.
  • Students. The name in a room comes from the token, with nothing to type. One person is one participant of the room, from any device. Per-address limits count the person: behind Teleport the whole class has one address, the proxy's.
  • Competitions. Entrants still sign in with their own key, as without the proxy.

Teleport. The application in Teleport points at Colloq's Service, not at an Ingress:

# teleport-kube-agent, app_service
apps:
  - name: colloq
    uri: http://colloq-app.colloq.svc:3000
    public_addr: colloq.teleport.corp.example

The chart values:

ingress:
  enabled: false                            # people come through Teleport
config:
  publicUrl: https://colloq.teleport.corp.example
  sso:
    jwksUrl: https://teleport.corp.example/.well-known/jwks.json
    issuer: teleport.corp.example           # the Teleport cluster name
    audience: http://colloq-app.colloq.svc:3000   # the app's uri: Teleport writes it into aud
    teacherRoles: colloq-teacher            # optional
networkPolicy:
  ingressController:
    from:                                   # the Teleport agent's Pods instead of ingress-nginx
      - namespaceSelector:
          matchLabels: {kubernetes.io/metadata.name: teleport}
        podSelector:
          matchLabels: {app: teleport-kube-agent}
  app:
    egress:
      - cidrs: [10.20.0.15/32]              # the Teleport proxy, for the keys
        ports: [443]

Instead of fetching the keys over the network, put the JWKS document in config.sso.jwks: the chart mounts it from the ConfigMap colloq-sso-jwks and no egress is needed, but a key rotation in Teleport needs a new value. The header is Teleport-Jwt-Assertion by default; config.sso.header sets another. Authorization won't do: the browser sends the room token in it, and a proxy writing its own there would break files, history and the Oracle. oauth2-proxy with --pass-access-token passes the token in X-Forwarded-Access-Token — fine when the IdP issues access tokens as JWTs (Keycloak does); the audience is then whatever the IdP writes into the access token.

Token fields. By default: who it is — sub or username; the name — name, traits.name, traits.full_name, traits.display_name, username; the email — email, traits.email, traits.mail, username; roles — roles, groups, traits.groups. config.sso.claims.* sets other paths, and the first non-empty one wins. If the token carries only a login, config.sso.emailDomain turns it into the address teachers are recognised by. A line at start says where the keys come from and whom Colloq adds; a refused token is logged once a minute with the reason.

Students without an account. Only someone with an account gets through Teleport. If students have none, keep the ordinary way in for them: an Ingress to the same Service at a public address and the class links, with config.publicUrl set to that public address, which links are built from. Teachers then open both the panel and the rooms at the Teleport address — the panel still builds the links for students from the public one: the token is checked the same way wherever the request came from, and a page at the Teleport address passes the origin check when the Teleport agent is in config.inbound.trustedProxies (the chart sets private).

What stays with the proxy. A token is a pass until it expires: if Colloq can also be reached around the proxy, a stolen token works until its exp. Keep the token lifetime short, and keep the token header out of access logs.

Storage

PVCWhat it holdsWho mounts itDefault size
colloq-dataThe SQLite database (colloq.db with its WAL journal), the link-signing key, the setup token, output images, class page files in page-files/, competition data, hidden answers and submissions, package sets, database snapshots in snapshots/the app; competition Pods — their own folders10 GiB; more for competitions with large data, and about 0.3–0.8 GiB for each 33-class course with class pages
colloq-workspaceRoom files, a <room>/ folder eachthe app; a room's Pod and its personal-notebook Pod — only that room's folder100 GiB

The storage class, size and access mode are set for both volumes at once (persistence.storageClass, persistence.accessMode) or for each (persistence.data.*, persistence.workspace.*); existingClaim attaches a PVC you already have. With persistence.retain: true, the default, the PVCs survive both helm uninstall and deleting the ArgoCD application.

ReadWriteMany: preferred. With RWX on both volumes, room Pods spread over the cluster wherever the scheduler puts them, and the installation's capacity is bounded by the quota rather than by one node. CephFS or NFSv4, where the protocol itself handles file locking, will do; SQLite has a single writer, the app. SQLite's own documentation warns that locking is unreliable on many NFS implementations: if in doubt, keep the database volume on block storage (persistence.data.accessMode: ReadWriteOnce), which brings back running every Pod with a volume on one node.

ReadWriteOnce: one node. This is the chart's default. An RWO volume attaches to one node, so everything that mounts Colloq's volumes must run there: if at least one of the volumes is RWO, the chart sets RUNTIME_COLOCATE_WITH_APP=1 for the broker, and every Pod that mounts a volume — rooms, personal notebooks, submissions, package preparation — gets a required podAffinity to the app's Pod (label colloq.dev/role: app, topology kubernetes.io/hostname). The app's own Pod prefers the node where rooms already run: after an upgrade the new Pod needs exactly the node the volume is attached to. The broker mounts no volume and can run anywhere. What follows: the whole installation's capacity is one node's free resources; a room beyond them doesn't start, and after 20 seconds of waiting the teacher sees what was short — memory, cores or a GPU; maintenance on that node is a break for all of Colloq. Pick a node with headroom; see Resources and quotas.

File ownership. Every Pod writes as uid and gid 1000. The app's Pod sets fsGroup: 1000, and a fresh block volume becomes writable for group 1000 by itself. For RWX volumes the CSI driver doesn't apply fsGroup by default (fsGroupPolicy: ReadWriteOnceWithFSType covers only RWO volumes with a filesystem type): the root of such a volume must be writable by uid 1000 on the storage side.

No per-room disk quota. The colloq-workspace volume is shared: a student who fills it stops file saving for everyone. Watch PVC usage — the kubelet metrics kubelet_volume_stats_used_bytes and kubelet_volume_stats_capacity_bytes, or colloq_disk_free_bytes from /metrics — with an alert at 80 %. If the StorageClass allows volume expansion, the PVCs can grow.

Images and the registry mirror

ImageWhat it is
ghcr.io/colloq-edu/colloq-app:vX.Y.ZThe app.
ghcr.io/colloq-edu/colloq-runtime:vX.Y.ZThe broker; also the submissions' result exporter and the package proxy.
ghcr.io/colloq-edu/colloq-kernel:vX.Y.Z-base, …-kaggle-baseRoom and submission kernels: the base and kaggle-base environments, CPU only.

Every image is built for amd64 only. A release and its chart share one version number. In the published chart the defaults carry the digest of every image of the release: the app's and the broker's in image.app.digest and image.runtime.digest, the kernels' in the environment catalog catalog.environments. The same values come as a file of their own, values-X.Y.Z.yaml, attached to the release on GitHub: handy for copying the images into a mirror and for diffing releases in a GitOps repository. The chart in the source tree carries no digests and refuses to render without them.

The environment chosen for a class is a catalog entry: a name, an image by digest, a GPU flag. An environment of your own — other libraries, or a GPU — is built outside the cluster from the repository's kernel/ directory with KERNEL_ENV=<name>, pushed to your registry and added to catalog.environments: on Kubernetes the panel doesn't build environments. The base-gpu image (about 9 GB of CUDA wheels) is not published; build a GPU environment the same way.

The registry mirror. global.imageRegistry replaces the registry in every image reference — the app's, the broker's, the catalog's kernels, the competition exporter and package proxy — and keeps the path and the digest. The value may carry a path too: a Harbor proxy cache of ghcr.io in a project called ghcr works as is, and harbor.corp.example/ghcr gives harbor.corp.example/ghcr/colloq-edu/colloq-app@sha256:…. So the mirror must hold the images under the same paths. Pull secrets go into global.imagePullSecrets: the app's and the broker's Pods get them, and every Pod the broker creates gets the first one.

Without internet access. On a machine that can reach ghcr.io:

# the release's chart
helm pull oci://ghcr.io/colloq-edu/charts/colloq --version X.Y.Z

# every image it refers to, the catalog's kernels included
helm template colloq colloq-X.Y.Z.tgz --set ingress.enabled=false \
  | grep -oE '[a-z0-9.-]+/[a-z0-9._/-]+@sha256:[a-f0-9]{64}' | sort -u

# each image into the mirror under the same path, with the same digest
crane copy ghcr.io/colloq-edu/colloq-app@sha256:… harbor.corp.example/colloq-edu/colloq-app:vX.Y.Z

# the chart into your OCI registry
helm push colloq-X.Y.Z.tgz oci://harbor.corp.example/charts

The chart package colloq-X.Y.Z.tgz is also attached to the release on GitHub, next to values-X.Y.Z.yaml. Any tool that copies the manifest as it is keeps the digest: crane copy, skopeo copy --all --preserve-digests or Harbor replication. Without internet access you lose notebook import from GitHub, the Oracle unless your own model runs on the internal network, and competition entrants' own packages unless you have a PyPI mirror. Everything else works: on Kubernetes students' code gets no internet anyway.

What the images are built from. The app and the broker stand on node:22-trixie-slim (Debian 13): the build applies every fix Debian has published so far, and npm, npx, corepack and yarn are removed from the image, since nothing runs them. Kernels stand on python:3.11-slim-bookworm (Debian 12). Every image in the registry carries an SBOM and provenance as BuildKit attestations: docker buildx imagetools inspect ghcr.io/colloq-edu/colloq-app:vX.Y.Z --format '{{json .SBOM}}'. Tools that copy the whole image index (crane copy, skopeo copy --all) carry them into the mirror with the image. On release day a scanner such as Trivy finds no critical vulnerability in the app and broker images. Kernels on Debian 12 still have critical findings Debian has not fixed (perl, zlib, sqlite); if your registry refuses to serve images with such findings, tell us: moving the kernels to Debian 13 is the next step.

Installation

Your team creates the namespace with its Pod Security labels and quota; then Helm or ArgoCD.

helm install colloq oci://ghcr.io/colloq-edu/charts/colloq --version X.Y.Z \
  --namespace colloq --values colloq-values.yaml

Values for a bank-like cluster: ReadWriteMany on CephFS, a Harbor proxy cache, secrets from Vault, an internal AI model and a PyPI mirror in Nexus, no internet.

global:
  imageRegistry: harbor.corp.example/ghcr       # a proxy cache of ghcr.io
  imagePullSecrets: [harbor-pull]

ingress:
  host: colloq.corp.example
  tls:
    secretName: colloq-tls                      # the organisation's certificate, type kubernetes.io/tls

persistence:
  accessMode: ReadWriteMany                     # without RWX keep ReadWriteOnce: everything on one node
  storageClass: cephfs

config:
  existingSecret: colloq-app-secrets            # OPENAI_API_KEY, SESSION_SECRET from Vault
  timezone: Europe/Moscow
  institution: Corporate University
  logFormat: json
  ai:
    provider: vllm
    baseUrl: https://llm.corp.example/v1
    model: corp-llm

dependencies:
  indexUrl: https://nexus.corp.example/repository/pypi/simple
  mirrorEgress:
    - cidrs: [10.20.0.15/32]                    # Nexus's addresses: a policy knows addresses, not names
      ports: [443]
  publicEgress: false

secrets:
  runtimeToken:
    existingSecret: colloq-internal             # key runtime-token
  roomSecret:
    existingSecret: colloq-internal             # key room-secret

rooms:
  maxMemory: 16Gi

networkPolicy:
  kubeApiServer:
    cidrs: [10.0.0.11/32, 10.0.0.12/32, 10.0.0.13/32]
    ports: [6443]
  app:
    egress:
      - cidrs: [10.30.0.20/32]                  # the internal AI model
        ports: [443]

metrics:
  enabled: true
  existingSecret: colloq-metrics                # key metrics-token
  serviceMonitor:
    enabled: true

Secrets from Vault come through External Secrets. In colloq-app-secrets the keys are the app's variable names (OPENAI_API_KEY, optionally SESSION_SECRET); leave config.ai.apiKey and config.sessionSecret empty then. colloq-metrics with the key metrics-token goes the same way as colloq-internal:

apiVersion: external-secrets.io/v1
kind: ExternalSecret
metadata:
  name: colloq-internal
  namespace: colloq
spec:
  refreshInterval: 1h
  secretStoreRef:
    kind: ClusterSecretStore
    name: vault
  target:
    name: colloq-internal
  data:
    - secretKey: runtime-token                  # openssl rand -hex 32, once
      remoteRef: {key: education/colloq, property: runtime-token}
    - secretKey: room-secret
      remoteRef: {key: education/colloq, property: room-secret}

ArgoCD. The chart is an OCI artifact; ArgoCD takes it as a Helm repository with enableOCI, its address without oci://.

apiVersion: v1
kind: Secret
metadata:
  name: corp-charts
  namespace: argocd
  labels:
    argocd.argoproj.io/secret-type: repository
stringData:
  type: helm
  name: corp-charts
  url: harbor.corp.example/charts
  enableOCI: "true"
---
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: colloq
  namespace: argocd
spec:
  project: education
  source:
    repoURL: harbor.corp.example/charts
    chart: colloq
    targetRevision: X.Y.Z
    helm:
      releaseName: colloq
      valuesObject:
        # the values above; or a values file from Git as a second source
        ingress:
          host: colloq.corp.example
  destination:
    server: https://kubernetes.default.svc
    namespace: colloq
  syncPolicy:
    automated:
      prune: true
      selfHeal: true

With persistence.retain: true (the default) the chart's PVCs carry argocd.argoproj.io/sync-options: Prune=false,Delete=false, so deleting the application in ArgoCD doesn't delete the data. Room and competition Pods are created by the broker and are not in Git. They deliberately carry no app.kubernetes.io/instance label: ArgoCD before 3.0 tracked its objects by it, and with prune would delete rooms in the middle of a class as superfluous. Your own labels for them go in rooms.podLabels; don't put app.kubernetes.io/instance there.

The owner link.

kubectl -n colloq exec deploy/colloq-app -- cat /data/setup-token

Open https://colloq.corp.example/admin/t/<token>, enter your name and email, and you are the owner. While the server has no owner, the same link is printed to the app's log at every start: kubectl -n colloq logs deploy/colloq-app | grep /admin/t/. Once there is an owner, the token is replaced and never printed again. The link is the key to the server: don't paste it into a group chat, and claim the server right after installing, before the log that carries it spreads through your log pipeline. Add teachers in the panel, under Who can teach: each gets a personal sign-in link.

Checks after installing.

  1. kubectl -n colloq get deploy,pods — colloq-app and colloq-runtime are Ready.
  2. curl -sS https://colloq.corp.example/api/readyz answers {"ok":true,"database":true,"workspace":true}.
  3. The full check, from inside the Pod, where error reasons are shown too: kubectl -n colloq exec deploy/colloq-app -- node -e 'fetch("http://127.0.0.1:3000/api/health").then(r => r.text()).then(console.log)'. The answer must have "ok":true and "isolation":"broker"; the capabilities field says whether competitions are ready.
  4. Two rooms from two devices on the students' network: shared editing, a cell run in each, a file upload. kubectl -n colloq get pods -l colloq.dev/role=kernel -o wide shows two Pods and their nodes.
  5. Isolation, from a cell in a room: connections to an outside address, to the app's Service and to the node's address must be refused.
    import socket
    for host, port in [("1.1.1.1", 443), ("colloq-app", 3000), ("<node address>", 10250)]:
        try:
            socket.create_connection((host, port), 3).close()
            print(host, port, "OPEN: check the network policies")
        except OSError as error:
            print(host, port, "closed:", error)
  6. If you use competitions: a trial submission through to the end; leaderboard rows should appear by themselves, with no page reload.
  7. Before the first group: a restore from a backup into a separate namespace.

Chart values

The main values and the variables they set. The full list, each value with an explanation, is the chart's values.yaml: helm show values oci://ghcr.io/colloq-edu/charts/colloq --version X.Y.Z. Other app settings from the Settings table of the VM guide go into config.extraEnv, or as keys of the Secret named in config.existingSecret.

ValueVariable and what it setsDefault
Address and ingress
ingress.hostThe hostname; required while the Ingress is on.—
config.publicUrlPUBLIC_URL — the external address links are built from.https://<ingress.host>
ingress.className, ingress.annotationsThe controller class and annotations.nginx, the annotations under Network
ingress.tls.secretName, ingress.tls.certManager.*The certificate's Secret, or a cert-manager issuer.colloq-tls
config.inbound.trustedProxiesTRUSTED_PROXIES — whose X-Forwarded-For to believe.private
config.inbound.sharedAddressesSHARED_ADDRESSES — NAT and VPN addresses that per-address caps don't apply to.empty
config.inbound.hstsHSTS — the app's Strict-Transport-Security header.true
Images and the catalog
global.imageRegistryA registry, optionally with a path, instead of ghcr.io for every image, the kernel catalog included.empty
global.imagePullSecretsSecrets for pulling images; broker Pods get the first one (RUNTIME_IMAGE_PULL_SECRET).—
image.app.digest, image.runtime.digestDigests of the app and broker images.the release's digests
catalog.environments, catalog.defaultEnvironmentKERNEL_CATALOG_FILE, RUNTIME_CATALOG_FILE — the environment catalog: a name, an image by digest, GPU; several revisions of one name with exactly one current: true.the release's base and kaggle-base
commonLabelsLabels on every chart object and every broker Pod (part of RUNTIME_POD_LABELS).—
Secrets
config.existingSecretAn existing Secret whose keys become the app's variables.—
secrets.runtimeToken.*, secrets.roomSecret.*The broker token and the room secret: existingSecret and existingSecretKey, or value.generated once by the chart (helm CLI only)
Storage
persistence.accessModeThe volumes' access mode. If either volume is ReadWriteOnce, the broker gets RUNTIME_COLOCATE_WITH_APP=1: every Pod with a volume runs on the app's node.ReadWriteOnce
persistence.storageClass, persistence.data.size, persistence.workspace.sizeThe storage class and volume sizes; existingClaim attaches a PVC you have.the cluster default; 10Gi and 100Gi
persistence.retainKeep the PVCs when the release or the ArgoCD application is deleted.true
Rooms
rooms.memoryRUNTIME_KERNEL_MEMORY — a room's memory unless the class sets its own.4Gi
rooms.maxMemoryRUNTIME_KERNEL_MEMORY_MAX — a room's memory ceiling. Set it explicitly: without it the broker takes its own node's memory minus 1 GiB.the broker node's memory − 1 GiB
rooms.cpuRUNTIME_KERNEL_CPU — a room's cores unless the class sets its own.2
rooms.ephemeralStorageRUNTIME_KERNEL_EPHEMERAL — space for a room's /tmp and home directory.2Gi
rooms.networkCOLLOQ_ROOM_NETWORK — with none, rooms get no DNS either.DNS only
rooms.nodeSelector, rooms.tolerationsRUNTIME_ROOM_NODE_SELECTOR, RUNTIME_ROOM_TOLERATIONS — where room and personal-notebook Pods run; competition Pods don't get them.—
rooms.priorityClassNameRUNTIME_PRIORITY_CLASS — the priority class of every Pod the broker creates.—
rooms.podLabels, rooms.podAnnotationsRUNTIME_POD_LABELS, RUNTIME_POD_ANNOTATIONS — labels and annotations on every broker Pod; the chart adds app.kubernetes.io/name, instance and part-of to the labels itself.—
runtime.inPlaceResizeRUNTIME_IN_PLACE_RESIZE — with false, live rooms aren't resized and the Role has no pods/resize rule.true
gpu.runtimeClassNameRUNTIME_GPU_RUNTIME_CLASS — the RuntimeClass of GPU rooms; empty means none.nvidia
gpu.nodeSelector, gpu.tolerationsRUNTIME_GPU_NODE_SELECTOR, RUNTIME_GPU_TOLERATIONS — added for GPU rooms.—
Network and outbound
networkPolicy.ingressController.fromWho is admitted to the app's port 3000.ingress-nginx
networkPolicy.dns.toWhere cluster DNS is.kube-dns in kube-system
networkPolicy.kubeApiServer.*The API servers' addresses and ports for the broker; ciliumEntity for Cilium.any address on 443 and 6443
networkPolicy.app.egressWhere the app may go outside: the model, a proxy, GitHub.nowhere
networkPolicy.monitoring.fromWho scrapes /metrics.Prometheus in monitoring
outbound.httpsProxy, outbound.httpProxy, outbound.noProxyHTTPS_PROXY, HTTP_PROXY, NO_PROXY — the app's outbound proxy. Service names (.svc, single-label names) and private addresses bypass it by themselves.—
outbound.extraCa.pem or .existingConfigMapNODE_EXTRA_CA_CERTS — the internal CA's certificates as PEM, for the app and package preparation.—
dependencies.indexUrl, dependencies.filesHostsDEPENDENCY_INDEX_URL, DEPENDENCY_FILES_HOSTS — the PyPI mirror for competition entrants' own packages: https, no credentials.public PyPI
dependencies.mirrorEgress, dependencies.publicEgressThe mirror's addresses and ports for the package proxy; the general rule for public addresses on 443.—; true
Classes and the Oracle
config.timezoneTZ — the time zone: day boundaries for daily quotas, dates.Europe/Moscow
config.uiLanguageUI_LANGUAGE — the interface language until the owner chooses one in the panel.ru
config.maxUploadMb, config.maxSessionMbMAX_UPLOAD_MB, MAX_SESSION_MB — the limit for one file and for all uploads to a room.50 and 1024
config.ai.provider, .baseUrl, .model, .apiKeyAI_PROVIDER, OPENAI_BASE_URL, OPENAI_MODEL, OPENAI_API_KEY — the Oracle's model; for a model on the internal network, vllm, ollama or custom. The key is better kept in config.existingSecret.openai, https://api.openai.com/v1, gpt-4o-mini
config.extraEnvOther app settings, as NAME: value.—
Monitoring, backups, resources
metrics.enabled, metrics.serviceMonitor.enabledMETRICS_TOKEN — /metrics behind a Bearer token, and a ServiceMonitor that uses the same token.off
config.logFormatLOG_FORMAT — with json, one JSON object per line.text
config.dbSnapshotHours, config.dbSnapshotKeepDB_SNAPSHOT_HOURS, DB_SNAPSHOT_KEEP — how often to write a consistent database snapshot into /data/snapshots, and how many to keep.24 h and 7
app.resources, runtime.resourcesRequests and limits of the app and the broker.see Resources and quotas

Resources and quotas

The numbers come from the limits in the code and from the same reasoning as for one VM, turned here into Pod requests and limits. The main difference from a VM: a room's Pod reserves its memory and cores in full rather than sharing them with its neighbours.

PodHow manyCPU: request / limitMemory: request / limitEphemeral storage
App1250m / 2512Mi / 2Gi—
Broker1100m / 1128Mi / 512Mi—
Roomone per running room2 / 24Gi / 4Gi2Gi
Personal notebooksone per room that uses themthe room's, or the class's own number, request equal to limit; about 250m per student is neededthe room's, or Memory per class; 1–1.5 GiB per student is needed2Gi
Submissionone per slotthe submission's cores + 100m / + 500mthe submission's memory + 64Mi / + 256Mi128Mi / 320Mi
Package preparationwhile it runs1100m / 1500m, proxy 100m / 500m2112Mi / 2304Mi, proxy 64Mi / 128Miup to 320Mi

The minimum quota for one class at a time (a room of 2 cores and 4 GiB), in the ResourceQuota keys requests.cpu, requests.memory, limits.cpu, limits.memory and pods:

ScenarioCPU requestsMemory requestsCPU limitsMemory limitsPods
(a) 30 students, one shared notebook2.354.6 GiB56.5 GiB3
(b) 30 students, each with a personal notebook (1.5 GiB and 0.25 core per student)9.8549.6 GiB12.551.5 GiB4
(c) 30 students and a competition, 7 slots of 2 cores and 2 GiB17.0519.1 GiB22.522.25 GiB10
(e) 100 people in a lecture room2.354.6 GiB56.5 GiB3

A room keeps its Pod while anyone is in it and for 2 hours after everyone has left, so back-to-back classes overlap: for each further class within two hours, add another room. Competition slots are one queue for the whole installation. A quota for scenario (c) with headroom for the next class's room and for package preparation:

apiVersion: v1
kind: ResourceQuota
metadata:
  name: colloq
  namespace: colloq
spec:
  hard:
    requests.cpu: "22"
    requests.memory: 28Gi
    limits.cpu: "30"
    limits.memory: 32Gi
    requests.ephemeral-storage: 10Gi
    pods: "25"
    services: "100"
    persistentvolumeclaims: "2"
    requests.storage: 150Gi
  • The API won't create a Pod over the quota, and the room or the submission gets a refusal. Keep the quota above the sum and the slot count within it: Executors at once in the Resources tab, or COMPETITION_SLOTS in config.extraEnv. The automatic count works from the cores and memory of the node the app runs on and knows nothing about your quota.
  • Services pile up: every room deleted for good leaves a headless Service without a ClusterIP, so give the Service count limit headroom.
  • A LimitRange with a per-container maximum must admit a room (2 cores and 4 GiB; a GPU room with its memory) and the largest submission you allow in competitions: 2 cores and 4 GiB by default.
  • With ReadWriteOnce the whole quota must fit on one node: for scenario (c), a node with at least 24 vCPU and 32 GiB free; for (b), at least 16 vCPU and 64 GiB.

The app, measured (30 Sep 2026, a 21 vCPU / 49 GB VM, nginx in front, make load): 500 people joined one room within a minute with no refusals, a median join taking 7 ms. The app process held 243 MB of memory and 20 % of one core at rest, and half a core with twenty people typing at once; an edit reached everyone else within 0.17 s (p95). At 300 people: 198 MB and the same shares of a core. This was measured on a VM, not in a cluster, but the process is the same: the chart's request of 512 MiB is twice the measured figure, and the limits of 2 cores and 2 GiB leave plenty of headroom.

Operations

Upgrades. helm upgrade colloq oci://ghcr.io/colloq-edu/charts/colloq --version X.Y.Z -n colloq -f colloq-values.yaml, or a new targetRevision in ArgoCD. The app and the broker update with the Recreate strategy: the old Pod stops before the new one starts, because SQLite has one writer. The break lasts seconds to a minute; a start with a database migration is given up to 10 minutes. Open rooms reconnect by themselves. Room Pods survive the upgrade: the broker leaves them alone when it stops, the new one adopts them, and Python variables stay. Placement settings — rooms.nodeSelector, rooms.tolerations, the priority class, the storage mode — go to a room's next Pod, while a running one stays where it is. Still, upgrade between classes.

Kernel revisions. A room stays on the kernel image revision it started with and runs only that one. On helm upgrade the chart keeps the previous release's revisions in the catalog with current: false by itself (at most catalog.maxRetainedPerEnvironment, three by default, per environment): it reads them from the live colloq-catalog ConfigMap, and rooms started earlier go on with their Python. ArgoCD renders without cluster access and cannot see the previous catalog: there a room whose revision is gone moves to its environment's current revision on its next start, and the app logs a line about it. To keep previous revisions under ArgoCD too, list them in catalog.environments with current: false. Keep the older revisions' images in the mirror while rooms use them.

The data version. If a new release raises the database schema version, the app first writes a snapshot, /data/snapshots/pre-schema-<from>-to-<to>-<time>.db; such snapshots are never removed automatically. An older release won't start on a database with a newer schema: the log says "colloq.db … was written by a newer Colloq", and the Pod goes into CrashLoopBackOff. So a rollback in that case is the older chart together with restoring the database: stop the app, put the snapshot in place of colloq.db from a maintenance Pod and delete colloq.db-wal and colloq.db-shm, then install the older chart version. Anything done after the upgrade is lost. COLLOQ_ALLOW_SCHEMA_DOWNGRADE=1 in config.extraEnv starts the older release on the newer database without a restore — only if you know the migration is harmless for it.

kubectl -n colloq scale deploy/colloq-app --replicas=0
kubectl -n colloq apply -f - <<'EOF'
apiVersion: v1
kind: Pod
metadata:
  name: colloq-maintenance
spec:
  securityContext:
    runAsNonRoot: true
    runAsUser: 1000
    runAsGroup: 1000
    seccompProfile: {type: RuntimeDefault}
  containers:
    - name: shell
      image: <the colloq-app image of the installed version>
      command: ["sleep", "3600"]
      securityContext:
        allowPrivilegeEscalation: false
        readOnlyRootFilesystem: true
        capabilities: {drop: ["ALL"]}
      resources:
        requests: {cpu: 100m, memory: 128Mi}
        limits: {cpu: 500m, memory: 256Mi}
      volumeMounts: [{name: data, mountPath: /data}]
  volumes:
    - name: data
      persistentVolumeClaim: {claimName: colloq-data}
EOF
kubectl -n colloq exec colloq-maintenance -- ls -l /data/snapshots/
kubectl -n colloq exec colloq-maintenance -- sh -c \
  'cp /data/snapshots/pre-schema-….db /data/colloq.db && rm -f /data/colloq.db-wal /data/colloq.db-shm'
kubectl -n colloq delete pod colloq-maintenance
helm rollback colloq <revision before the upgrade> -n colloq

Backups. Snapshot both PVCs with your own tools — Velero with CSI snapshots, or your storage's snapshots — and the namespace's Secrets unless they come from Vault, in which case Vault is their backup. A snapshot of a live SQLite volume is a crash image: the database usually recovers from it, but it is not something to rely on. So every config.dbSnapshotHours hours (24 in the chart) the app itself writes a consistent snapshot, /data/snapshots/colloq-<time>.db, and keeps the last config.dbSnapshotKeep (7): every volume snapshot holds a whole copy of the database no older than a day. The volumes are snapshotted separately, so room files and the database in a backup can be minutes apart. Copies of the colloq-data volume contain the teachers' sign-in keys, the setup token and the model key if it was entered in the panel: keep them as secrets. Test a restore into a separate namespace at least once a term. The colloq-host backup and restore commands are for the VM and are not used in a cluster.

Monitoring. The liveness probe is /api/livez, readiness is /api/readyz: it checks only the database and the working files, so a broker or API hiccup doesn't take the site out of rotation, and kernel trouble shows where code is run and in /api/health. With metrics.enabled, /metrics serves Prometheus text to requests carrying Authorization: Bearer <token>; without the setting it answers 404. Useful for alerts: free space on the volumes (colloq_disk_free_bytes), the time of the last database snapshot (colloq_db_snapshot_last_timestamp_seconds), notebook save and kernel start failures (colloq_notebook_save_failures_total, colloq_kernel_start_failures_total). Logs go to stdout and stderr; with config.logFormat: json each line is a JSON object with time, level, msg and context fields. The log contains room addresses, and a room address is a sign-in link, so access to the log equals access to the rooms. Staff can use GET /api/instance/operations: queue age, retries, notebook save failures, memory reservations, free disk.

GPU nodes. A GPU room requests nvidia.com/gpu: 1 and the RuntimeClass from gpu.runtimeClassName (nvidia by default; empty means none); gpu.nodeSelector and gpu.tolerations pick the nodes. The driver, the device plugin or GPU Operator and the RuntimeClass itself are cluster-scoped objects your team provides. One card per room, not shared, unless your device plugin splits cards itself. Personal notebooks and competition submissions never get a GPU. Build the GPU environment's image yourself; see Images and the registry mirror.

Node maintenance. Room Pods are bare Pods without a controller and with emptyDir: kubectl drain removes them only with --force --delete-emptydir-data, and the cluster autoscaler won't free a node that runs them until you add the annotation cluster-autoscaler.kubernetes.io/safe-to-evict: "true" through rooms.podAnnotations. A room brings a deleted Pod back at its next run: notebooks and files are there, Python variables are not. With ReadWriteOnce, maintenance on the app's node is a break for the whole installation.

Uninstalling. helm uninstall removes the chart's objects except the PVCs — with persistence.retain: true they stay — and leaves the Pods and Services the broker created: they don't belong to Helm. Delete them by label: kubectl -n colloq delete pod,svc -l app.kubernetes.io/managed-by=colloq-runtime.

Honest limits

  • One app replica. The database is SQLite with one writer: there is no high availability, and an upgrade or a node failure interrupts work while the Pod restarts. To grow, run separate installations, one per namespace — for example one per department.
  • SSO only through a proxy. Colloq has no SSO, LDAP or second-factor sign-in of its own: a proxy that signs a JWT provides it — see Sign-in through your SSO. Without such a proxy the owner and teachers sign in with personal links, and the panel can be closed off at the ingress controller — for example oauth2-proxy with the corporate IdP through ingress-nginx's external authentication annotations, on a separate Ingress for the paths /admin and /api/admin. Competition entrants sign in with their own key even behind the proxy.
  • A room link is a shared link. Whoever has it can join and, by default, run code; it can't be revoked. Rooms don't suit individually graded or confidential work: participants share the kernel, the files and the Linux user.
  • ReadWriteOnce means one node. The whole installation fits on one node, and its maintenance is a break for everyone.
  • Students' code has no internet. pip install in a cell doesn't work; libraries come only in the catalog's kernel images, built outside the cluster.
  • No per-room disk quota: the colloq-workspace volume is shared.
  • Network policies are not all of the isolation. Standard NetworkPolicy doesn't guarantee that a Pod can't reach services on its own node, and Pods share the node's Linux kernel.
  • amd64 only.
  • Kernels on Debian 12. A scanner finds critical findings in them that Debian has not fixed; see “Images and the registry mirror”.
  • A new path. The chart is new in this release: test the installation in a non-production namespace before running classes in it.