Skip to content

Latest commit

 

History

History
160 lines (112 loc) · 5.81 KB

File metadata and controls

160 lines (112 loc) · 5.81 KB

Helpful procedures

Checking server health

These examples are written with reference to ed0 but apply to other instances running the full stack.

First check the status page. If this page is being served, then the instance is running and does not need to be rebooted at this stage.

If ALL of the experiments have streams missing, then it is likely that the relay requires restarting. To restart the relay, log in to the server

cd ~/sources/admin-tools/ed0
./login.sh

Once logged in, check the current load using top

top

Typically you should see around 40% load on relay, and about 20% load on nginx at rest, and under heavy usage the relay load approaches 70%. If there is no load then there is either a problem with the relay, or with the experiments themselves (e.g. campus network outage).

Exit top with q

Check the status of the relay service

sudo systemctl status relay

If this is active but there is NO load showing in top, then it needs restarting. If it is inactive then it needs restarting. Restart it:

sudo systemctl restart relay

This can take a minute or two. If it takes longer than two minutes, then you can control-C and try

sudo systemctl status relay 
sudo system stop relay #if it is still running, then try stopping it
sudo system start relay #else just try starting it

but really, it amounts to the same thing. After restarting, check the service is active

sudo systemctl status relay

And if it is active, check the load on it in top

top

If these steps have not been successful, then check the disk capacity

sudo df ./ -h

If there is more than 50% disk used, then it's likely there is a build up of log files. Currently, status is running in debug mode, producing GB of logging. You can safely delete the status server logs at this time. This should return the disk usage to closer to 31% (although df only reports the correct value after a delay/status is restarted).

cd /var/log/status
ls -alh 
sudo rm status.log
sudo systemctl restart status

Don't worry if the first few re-freshes of the status page appear to show the experiments are down - it takes a few minutes for status to collect information and start reporting the correct state of the system, then a further period of STATUS_HEALTH_STARTUP=15m until status will start adjusting the availability of experiments. Experiments will show as unavailable on the status, but bookable anyway.

If the disk is too full, you won't be able to log on, so if in doubt, better to delete the status.log at this time.

If the instance needs restarting, then go to the google cloud console, choose the practable.io organisation, then the app-practable-io-alpha project, then Compute Engine in the menu on the left, and then select the instance, and restart it (see option under the three vertical dots in the right hand column of the entry for the instance in the list of VMs).

Note that restarting the instance has the following consequences

  • timeout while system restarts
  • all user bookings are lost (if the instance is still running, you can export and re-import after restart, but if it is not running, there is no way to export and you must restart)
  • the booking manifest needs re-uploading

To upload the manifest:

cd ~/sources/manifest-ed0
./check.sh
./upload.sh

Disk full issues

If you are trying to identify the source of disk usage, then the du tool is helpful.

sudo du -h -d 1

Snaps sometimes fill up ... but the minimum number of versions that must be retained is two.

Consider editing the terraform plan and increasing the disk size.

Adding a new user interface to the ed0

Assuming it is working on the ed-dev-ui server, then:

  • rebuild the UI for the ed0 basepath
  • add to the static repo for ed0 and git push to run webhook (to be implemented)
  • add an entry in the ./ed0/templates/nginx.conf.template for the new UI
  • run the ed0 ./configure.sh script
  • run the ansible playbook ansible-playbook update-nginx-conf.yaml
  • edit the manifest in github.com/practable/manifest-ed0

Serving WebAssembly (.wasm) assets

UIs that ship WebAssembly (e.g. verisim-1.0) must be served with Content-Type: application/wasm, otherwise the browser refuses to stream-compile the module (WebAssembly.instantiateStreaming fails and the UI does not start).

The nginx packaged with Ubuntu 22.04 (nginx 1.18) predates the upstream addition of this mapping, so /etc/nginx/mime.types has no wasm entry:

grep -n 'wasm' /etc/nginx/mime.types   # no output on an unpatched server

The mapping is managed by ansible, in both install-nginx (new servers) and update-nginx-conf (existing servers), so it survives re-provisioning and subsequent configuration runs. Because nginx.conf already does include /etc/nginx/mime.types;, no change to nginx.conf.template is needed.

To apply it to an existing server, from the instance directory (e.g. ed0-alternate):

cd ~/sources/admin-tools/ed0-alternate
./configure.sh
cd playbooks
ansible-playbook update-nginx-conf.yml

The playbook validates the configuration with nginx -t before the handler reloads nginx, and only reloads if something actually changed. A second run reports changed=0 and does not reload.

Verify on the server:

grep -n 'wasm' /etc/nginx/mime.types
# 55:    application/wasm                                 wasm;

And over HTTP, once the UI is deployed (ed0 and ed0-alternate both serve the ed0 instance path on app.practable.io):

curl -sSI https://app.practable.io/ed0/static/ui/verisim-1.0/<asset>.wasm \
  | grep -Ei '^HTTP|^Content-Type'
# HTTP/2 200
# content-type: application/wasm

Substitute the instance path for other servers, e.g. ed0-staging, ed-dev-ui, or https://test.practable.io/ed1/... for ed1-test.