configpolicy

Author	SHA1	Message	Date
Dustin C. Hatch	53b39338dd	r/postgresql-server: Add script to upgrade database The `postgresql-upgrade.sh` script arranges to run `pg_upgrade` after a major PostgreSQL version update. It's scheduled by a systemd unit, _postgresql-upgrade.service_, which runs only after an OS update.	2024-11-17 10:27:31 -06:00
Dustin C. Hatch	0048a87630	r/postgresql-server: Set become on postgres tasks Tasks that must run as the _postgres_ user need to explicity enable `become`, in case it is not already enabled at the playbook level. This can happen, for example, when the playbook is running directly as root.	2024-11-16 11:50:28 -06:00
Dustin C. Hatch	2d5f9e66c1	chromie: Scrape logs from serial consoles Now that we have the serial terminal server managing `picocom` processes for each serial port, and those `picocom` processes are configured to log console output to files, we can configure Promtail to scrape these log files and send them to Loki.	2024-11-10 18:34:49 -06:00
Dustin C. Hatch	a82700a257	chromie: Configure serial terminal server	2024-11-10 13:15:08 -06:00
Dustin C. Hatch	6115762847	r/serterm: Deploy serial terminal multiplexer Using `tmux`, we can spawn a bunch of `picocom` processes for the serial ports connected to other server's console ports. The _serial-terminal-server_ service manages the `tmux` server process, while the individual _serial-terminal-server-window@.service_ units create a window in the `tmux` session. The serial terminal server runs as a dedicated user. The SSH server is configured to force this user to connect to the `tmux` session. This should help ensure the serial consoles are accessible, even if the Active Directory server is unavailable.	2024-11-10 13:15:08 -06:00
Dustin C. Hatch	8b9cf1985a	r/wal-g-pg: Schedule weekly delete jobs WAL-G slows down significantly when too many backups are kept. We need to periodically clean up old backups to maintain a reasonable level of performance, and also keep from wasting space with useless old backups.	2024-11-05 19:28:57 -06:00
Dustin C. Hatch	eaf9cbef9a	Merge remote-tracking branch 'origin/frigate-exporter'	2024-11-05 07:01:31 -06:00
Dustin C. Hatch	c1dc52ac29	Merge branch 'loki'	2024-11-05 07:01:13 -06:00
Dustin C. Hatch	39d9985fbd	r/loki-caddy: Caddy reverse proxy for Loki Caddy handles TLS termination for Loki, automatically requesting and renewing its certificate via ACME.	2024-11-05 06:54:27 -06:00
Dustin C. Hatch	010f652060	hosts: Add loki1.p.b _loki1.pyrocufflink.blue_ replaces _loki0.pyrocufflink.blue_. The former runs Fedora Linux and is managed by Ansible, while the latter ran Fedora CoreOS and was managed by Ignition and _cfg_.	2024-11-05 06:54:27 -06:00
Dustin C. Hatch	abfd35a68e	raid-array: Create udev rules to auto re-add disks This udev rule will automatically re-add disks to the RAID array when they are connected. `mdadm --udev-rules` is supposed to be able to generate such a rule based on the `POLICY` definitions in `/etc/mdadm.conf`, but I was not able to get that to work; it always printed an empty rule file, no matter what I put in `mdadm.conf`.	2024-11-05 06:52:20 -06:00
Dustin C. Hatch	168bfee911	r/webites: Add apps.du5t1n.xyz F-Droid repo I want to publish the _20125_ Status application to an F-Droid repository to make it easy for Tabitha to install and update. F-Droid repositories are similar to other package repositories: a collection of packages and some metadata files. Although there is a fully-fledged server-side software package that can manage F-Droid repositories, it's not required: the metadata files can be pre-generated and then hosted by a static web server just fine. This commit adds configuration for the web server and reverse proxy to host the F-Droid repository at _apps.du5t1n.xyz_.	2024-11-05 06:47:02 -06:00
Dustin C. Hatch	7e8aee072e	r/bitwarden_rs: Redirect to canonical host name Bitwarden has not worked correctly for clients using the non-canonical domain name (i.e. _bitwarden.pyrocufflink.blue_) for quite some time. This still trips me up occasionally, though, so hopefully adding a server-side redirect will help. Eventually, I'll probably remove the non-canonical name entirely.	2024-11-05 06:37:03 -06:00
Dustin C. Hatch	0807afde57	r/dch-proxy: Use separate sockets for HTTP v4/v6 Although listening on only an IPv6 socket works fine for the HTTP front-end, it results in HAProxy logging client requests as IPv4-mapped IPv6 addresses. For visual processing, this is ok, but it breaks Loki's `ip` filter.	2024-11-05 06:34:55 -06:00
Dustin C. Hatch	90351ce59e	r/dch-proxy: Include host name in log messages When troubleshooting configuration or connection issues, it will be helpful to have the value of the HTTP Host header present in log messages emitted by HAProxy. This will help reason about HAProxy's routing decisions. For TLS connections, of course, we don't have access to the Host header, but we can use the value of the TLS SNI field. Note that the requisite `content set-var` directive MUST come before the `content accept`; HAProxy stops processing all `tcp-request content ...` directives once it has encountered a decision.	2024-11-05 06:32:49 -06:00
Dustin C. Hatch	370a1df7ac	dch-proxy: Proxy for dynk8s-provisioner The reverse proxy needs to handle traffic for the _dynk8s-provisioner_ in order for the ephemeral Jenkins worker nodes in the cloud to work properly.	2024-11-05 06:30:02 -06:00
Dustin C. Hatch	3ca94d2bf4	r/haproxy: Enable Prometheus metrics HAProxy can export stats in Prometheus format, but this requires special configuration of a dedicated front-end. To support this, the _haproxy_ Ansible role now has a pair of variables, `haproxy_enable_stats` and `haproxy_stats_port`, which control whether or not the stats front-end is enabled, and if so, what port it listens on. Note that on Fedora with the default SELinux policy, the port must be labelled either `http_port_t` or `http_cache_port_t`.	2024-11-05 06:23:49 -06:00
Dustin C. Hatch	9f30998fbf	r/jellyfin: Enable Prometheus metrics Jellyfin can expose metrics in Prometheus format, but this functionality is disabled by default. To enable it, we must set `EnableMetrics` in the configuration file. This commit adds a template configuration file that uses the `jellyfin_enable_metrics` Ansible variable to control this value.	2024-11-05 06:21:38 -06:00
Dustin C. Hatch	a9923dcb57	hosts: chromie: Enable collectd md, thermal plugins To monitor the RAID array and various temperature probes.	2024-11-04 17:52:46 -06:00
Dustin C. Hatch	29d65dd0d5	gw1: squid: Allow access to Gitea Specifically to allow _nvr2.pyrocufflink.blue_ to fetch the _frigate-exporter_ container image.	2024-10-21 20:27:31 -05:00
Dustin C. Hatch	ccf33f90e0	r/frigate-exporter: Deploy Prometheus exporter Frigate exports useful statistics natively, but in a custom JSON format. There is a [feature request][0] to add support for Prometheus format, but it's mostly being ignored. A community member has created a standalone process that converts the JSON format into Prometheus format, though, which we can use. [0]: https://github.com/blakeblackshear/frigate/issues/2266	2024-10-21 20:27:31 -05:00
Dustin C. Hatch	4cd983d5f4	loki: Add role+playbook for Grafana Loki The current Grafana Loki server, loki0.pyrocufflink.blue, runs Fedora CoreOS and is managed by Ignition and cfg. Since I have declared cfg a failed experiment, I'm going to re-deploy Loki on a new VM running Fedora Linux and managed by Ansible. The loki role installs Podman and defines a systemd-managed container to run Grafana Loki.	2024-10-20 12:10:55 -05:00
Dustin C. Hatch	4ac79ba18d	minio-backups: No syslog for nginx access logs MinIO/S3 clients generate a _lot_ of requests. It's also not particularly useful to have these stored in Loki anyway. As such, we'll stop routing them to syslog/journal. Having access logs is somewhat useful for troubleshooting, but really for only live requests (i.e. what's happening right now). We therefore keep the access logs around in a file, but only for one day, so as not to fill up the filesystem with logs we'll never see.	2024-10-20 12:10:17 -05:00
Dustin C. Hatch	388fd91096	r/nginx: Configure error/access syslog separately There may be cases where we want either error logs or access logs to be sent to syslog, but not both. To support these, there are now two variables: `nginx_access_log_syslog` and `nginx_error_log_syslog`. Both use the value of the `nginx_log_syslog` variable by default, so existing users of the _nginx_ role will continue to work as before.	2024-10-20 12:10:17 -05:00
Dustin C. Hatch	4ae25192d0	vm-hosts: Fix domain label The `__path__` label is automatically changed to `filename` before the processing pipeline begins.	2024-10-14 12:32:25 -05:00
Dustin C. Hatch	36145cb2ee	minio-backups: Disable nginx log files We don't need local log files when messages are already stored locally in the journal and remotely in Loki.	2024-10-14 12:00:19 -05:00
Dustin C. Hatch	845911dcbd	r/nginx: Make logging to files optional If _nginx_ is configured to send error/access log messages to syslog, it may not make sense to _also_ send messages to log files as well. The `nginx_error_log_file` and `nginx_access_log_file` variables are now available to control whether/where to send log messages. Setting either of these to a falsy value will disable logging to a file. A non-empty string value is interpreted as the path to a log file. By default, the existing behavior of logging to `/var/log/nginx/error.log` and `/var/log/nginx/access.log` is preserved.	2024-10-14 12:00:19 -05:00
Dustin C. Hatch	a0c5ffc869	postgresql: Collect Wal-G metrics with statsd_exporter _wal-g_ can send StatsD metrics when it completes an upload/backup/etc. task. Using the `statsd_exporter`, we can capture these metrics and make them available to Victoria Metrics.	2024-10-13 20:01:19 -05:00
Dustin C. Hatch	87b9014721	r/statsd-exporter: Deploy statsd exporter The statsd exporter is a Prometheus exporter that converts statistics from StatsD format into Prometheus metrics. It is generally useful as a bridge between processes that emit event-based statistics, turning them into Prometheus counters and gauges.	2024-10-13 19:59:52 -05:00
Dustin C. Hatch	a22c8aa0d2	r/nextcloud: Configure trashbin retention Setting the `trashbin_retention_obligation` setting to `auto, 30` should supposedly delete files in users' trash bins after 30 days.	2024-10-13 18:38:12 -05:00
Dustin C. Hatch	265aa074aa	r/nextcloud: Configure Memories app The [Memories] app for Nextcloud provides a better user interface and more features than the built-in Photos app. The latter seems to be somewhat broken recently (timeline stops in June 2024, even though there are more recent photos available), so we're trying out Memories (and Recognize for facial recognition). [Memories]: https://memories.gallery	2024-10-13 18:36:25 -05:00
Dustin C. Hatch	5ab0bcd5bf	r/nextcloud: Update rewrite config for .mjs files Nextcloud 28+ uses JavaScript modules (`.mjs` files). These need to be served from the filesystem like other static files, so the mod_rewrite configuration needs to be updated as such.	2024-10-13 18:35:01 -05:00
Dustin C. Hatch	221d3a2be9	vm-hosts: Scrape libvirt logs with Promtail Collecting logs from VM serial consoles and QEMU monitor.	2024-10-13 18:33:25 -05:00
Dustin C. Hatch	1e6ab546bc	r/vmhost: Create directory for console logs Need a directory where _libvirt_ can write logs from VM serial console output.	2024-10-13 18:30:04 -05:00
Dustin C. Hatch	75a146e19e	newvm: Configure serial console log file When a VM uses a serial port for its default console, kernel messages (e.g. panics) are lost if no console client is connected at the time. This is a major disadvantage when compared to a graphical console, which usually at least keeps a "screenshot" of the console when the kernel crashes. While researching the available console device types to determine how best to implement a tool that would both log the output from the serial console at all times, while still allowing interactive connections to it, I discovered that _libvirt_ actually already has this exact functionality built-in: https://libvirt.org/formatdomain.html#consoles-serial-parallel-channel-devices	2024-10-13 18:12:46 -05:00
Dustin C. Hatch	9bea8e1ce7	nextcloud: Scrape logs with Promtail Nextcloud writes JSON-structured logs to `/var/lib/nextcloud/data/nextcloud.log`. These logs contain errors, etc. from the Nextcloud server, which are useful for troubleshooting. Having them in Loki will allow us to view them in Grafan as well as generate alerts for certain events.	2024-10-13 18:05:50 -05:00
Dustin C. Hatch	ceaef3f816	hosts: Decommission burp1.p.b Everything has finally been moved to Chromie.	2024-10-13 17:52:48 -05:00
Dustin C. Hatch	808a912630	websites: Remove proxy roles Reverse proxy for web sites and applications accessible to the Internet is now handled by HAProxy.	2024-10-13 12:54:50 -05:00
Dustin C. Hatch	5ced24f2be	hosts: Decommission matrix0.p.b The Synapse server hasn't been working for a while, but we don't use it for anything any more anyway.	2024-10-13 12:53:49 -05:00
Dustin C. Hatch	219fe75424	r/nginx: logrotate: do not delay compressing _nginx_ access logs are typically either very small or very large. For small log files, it's fast enough to decompress them on the fly if necessary. For large files, they may take up so much space in uncompressed form that the log volume fills too quickly. In either case, compressing the files as soon as they are rotated is a good option, especially since their contents should already be sent to Loki.	2024-09-30 12:43:25 -05:00
Dustin C. Hatch	dfdddd551f	minio-backups: Keep nginx logs for 3 days _WAL-G_ and _restic_ both generate a lot of HTTP traffic, which fills up the log volume pretty quickly. Let's reduce the number of days logs are kept on the file system. Logs are shipped to Loki anyway, so there's not much need to have them local very long.	2024-09-29 11:21:24 -05:00
Dustin C. Hatch	829c04332d	r/nginx: Configure logrotate The default `logrotate` configuration for _nginx_ may not be appropriate for high-volume servers. The `nginx_keep_num_logs` variable is now available to control how many days of logs are kept.	2024-09-29 11:20:29 -05:00
Dustin C. Hatch	0353360360	dch-proxy: Allow Internet access to IN Invoice Ninja needs to be accessible from the Internet in order to receive webhooks from Stripe. Additionally, Apple Pay requires contacting Invoice Ninja for domain verification.	2024-09-10 12:01:00 -05:00
Dustin C. Hatch	9e610eaf11	r/minio-backups-cert: Enable/start cerbot timer Forgot to ensure the _certbot-renew.timer_ unit was enabled and started, so the MinIO certificate did not get renewed the first time.	2024-09-08 09:15:36 -05:00
Dustin C. Hatch	621f82c88d	hosts: Migrate remaining hosts to Restic Gitea and Vaultwarden both have SQLite databases. We'll need to add some logic to ensure these are in a consistent state before beginning the backup. Fortunately, neither of them are very busy databases, so the likelihood of an issue is pretty low. It's definitely more important to get backups going again sooner, and we can deal with that later.	2024-09-07 20:45:24 -05:00
Dustin C. Hatch	7d93ba836e	r/restic: Enhance restic-backup security sandbox Since `restic` needs to run as root in order to back up files regardless of their permissions, we need to restrict it to doing only that. Using systemd sandbox features, especially the capability bounding set, we can remove all of _root_'s powers except the ability to read all files.	2024-09-04 17:43:24 -05:00
Dustin C. Hatch	c2c283c431	nextcloud: Back up Nextcloud with Restic Now that the database is hosted externally, we don't have to worry about backing it up specifically. Restic only backs up the data on the filesystem.	2024-09-04 17:41:42 -05:00
Dustin C. Hatch	0f4dea9007	restic: Add role+playbook for Restic backups The `restic.yml` playbook applies the _restic_ role to hosts in the _restic_ group. The _restic_ role installs `restic` and creates a systemd timer and service unit to run `restic backup` every day. Restic doesn't really have a configuration file; all its settings are controlled either by environment variables or command-line options. Some options, such as the list of files to include in or exclude from backups, take paths to files containing the values. We can make use of these to provide some configurability via Ansible variables. The `restic_env` variable is a map of environment variables and values to set for `restic`. The `restic_include` and `restic_exclude` variables are lists of paths/patterns to include and exclude, respectively. Finally, the `restic_password` variable contains the password to decrypt the repository contents. The password is written to a file and exposed to the _restic-backup.service_ unit using [systemd credentials][0]. When using S3 or a compatible service for respository storage, Restic of course needs authentication credentials. These can be set using the `restic_aws_credentials` variable. If this variable is defined, it should be a map containing the`aws_access_key_id` and `aws_secret_access_key` keys, which will be written to an AWS shared credentials file. This file is then exposed to the _restic-backup.service_ unit using [systemd credentials][0]. [0]: https://systemd.io/CREDENTIALS/	2024-09-04 09:40:29 -05:00
Dustin C. Hatch	708bcbc87e	Merge remote-tracking branch 'refs/remotes/origin/master'	2024-09-03 17:18:18 -05:00
Dustin C. Hatch	dce7908a94	chromie: Set MinIO root password	2024-09-02 21:24:59 -05:00

1 2 3 4 5 ...

997 Commits