Services
MUSE
Multi-server Unified Service Environment
A management system that shows resources, jobs and per-user usage across the lab's compute servers and cluster on one screen.
What it does
MUSE is a management system for observing the lab's compute servers and cluster in one place. It shows in real time how much CPU, GPU, memory and storage each server is using, who is using how much, and what is sitting in the cluster queue.
Use it to check which servers are free right now before you run a computation. The lab's analysis services also look at these observations before submitting heavy computation.
MUSE only observes. You do not submit jobs or reserve resources from the dashboard — reservation is done by the cluster scheduler.
What to know before you start
You log in with your lab server account. Not the contextBio unified account, but the account you use on the lab's compute servers. Even with a server account, you must be on the access allow-list to get in. If you cannot log in, ask the lab administrator to allow you.
It is a system outside this site. Pressing 서비스 바로가기 (Go to service) on the MUSE card in
the "Management systems" group of the company site's service list opens MUSE in a new tab.
There is no separate MUSE screen inside the site.
It is not the similarly named tool. MuSE in AURORA's list of variant callers is a somatic variant-calling program and has nothing to do with this system.
How to use it
- Press
서비스 바로가기(Go to service) on the MUSE card in the site's service list - Log in with your lab server account
- Look at the summary cards and the issue history at the top of the dashboard to see whether any server has a problem right now
- Look at the cluster queue in the Slurm panel. Pending jobs show the reason they are waiting — this is where to check "why isn't my job running"
- Use the trend charts and tables to see per-server and per-user usage, and decide where to put your computation
What you see on screen
The dashboard is laid out top to bottom like this.
| Position | What it shows |
|---|---|
| Summary cards | Overall status |
| Issue history | Anomalies such as stopped collection, low disk or memory, GPU throttling, and their history. Urgent ones appear as a red band under the header |
| Slurm | Running and pending jobs in the current queue, and node states |
| Playback bar and range selector | Wind the trend charts below back to an earlier point in time |
| Trend charts | Top 5 users by CPU, memory and GPU; CPU and GPU utilisation per server |
| Tables | Per-user usage, per-server status, server specifications |
Beyond the dashboard there are a few more pages — the list of programs installed on the servers, the data inventory of shared storage, the server operating rules, and an operating playbook.
Must-know points
The numbers in the Slurm panel are reservations, not measurements. They show the CPU and GPU a job requested; for what is actually being used, look at the measured values in the server table. It is normal for the two to disagree.
There is no record of finished jobs. The Slurm panel is a snapshot of the current queue. A finished job drops off the list within a few minutes and cannot be seen again. Very short jobs may pass between refreshes and never appear at all.
No observation means "unknown", not "free". An empty cell for a server whose collection has stopped does not mean that server is idle. Check whether the issue history shows stopped collection.
The cluster does not isolate GPUs. A job that did not request a GPU can still see the node's GPUs. Use the measured values in the server table to check whether your GPU overlaps with someone else's job.
Only administrators stop processes. The administrator screen has a fleet-wide process list and stop controls, but it is not open to ordinary users. If a job needs to be stopped, ask an administrator.
Software availability
This page has not verified MUSE's full installation. Only the confirmed components are listed.
| Component | What it is used for | Symptom when missing or disconnected |
|---|---|---|
| Per-server collection agent | Collecting each server's CPU, GPU, memory, storage and per-user usage | That server shows up in the issue history as collection stopped |
| Slurm query | Queue and node states | The Slurm panel changes to "controller not responding" (red) or "not deployed" (grey) |
| Time-series store | Trend charts | An update-failed line appears under the trend charts. Trends go back at most 15 days |
What to do when blocked
- Read the exact text shown on the panel — an empty table, "not responding" and "update failed" are different situations
- If the whole screen will not open or login is blocked, tell the lab administrator
- If only one server shows stopped collection, check its state with the administrator before putting new computation on it
Last updated ·
Docs