Tubklima

Tubklima is the cluster and Linux operations environment at TU Berlin’s climatology group. It appears here because most of the infrastructure notes in this garden assume a real working context, and this is it.

That context shapes what those notes say. Advice about HPC written from a specification sheet differs from advice written by someone who has waited in a queue, filled a filesystem, and had to explain to colleagues why their jobs stopped. The storage notes — BeeGFS, Lustre, Ceph — compare designs on operational burden as much as on throughput, because operational burden is what actually decides these things at institutional scale.

What the environment is like

It is the mid-sized institutional case rather than either extreme: substantially more than a workstation, far from a national supercomputing centre. That middle is where most research computing actually happens and it gets written about least.

The characteristic constraint is that nobody’s full-time job is any one of these systems. There is no dedicated storage engineer, no dedicated security team, no dedicated monitoring team. The same people administer the cluster, run the models, and answer questions — which is why a tool being simple to operate correctly beats a tool being more capable, and why that consideration recurs throughout the notes here.

Multi-machine automation through SaltStack and operational visibility through Checkmk both matter more in this setting than they would with more staff, because they substitute for attention nobody has to spare.

The climate work runs here: WRF for downscaling, the Galapagos analysis, and CER v2 over Berlin-Brandenburg.

Specific hostnames, network layout, and access arrangements stay in private notes. The general shape is what makes the operational notes legible.

See also: HPC, Linux administration in datacenters, and the computing map.