Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions README.es.md
Original file line number Diff line number Diff line change
Expand Up @@ -367,6 +367,10 @@ Dependencias Go (mantenidas a 2 a propósito):

## Registro de cambios

### v2.9.20

- **Reduce el volumen de logs sudo/journald de los colectores periódicos (#124)**: el servicio de EasyZFS generaba decenas de miles de entradas sudo al día porque cada comando `zpool`/`zfs` que ejecutaban los colectores se logueaba en `sudo` como comando permitido. Se añade `Defaults:easyzfs !pam_session, !log_allowed` al sudoers para que los comandos permitidos no se logueen; los intentos fallidos siguen registrándose. Se complementa con cachés TTL en `ZpoolCollector` para propiedades estables del pool (`autotrim`, `checkpoint`, `compressratio`) y para `zpool status -t`, y una nueva variable de entorno `EASYZFS_ZPOOL_INTERVAL` (60 s por defecto) para controlar el intervalo del colector. Validado en nodos reales de Proxmox: `journalctl -g 'easyzfs :'` pasa de miles de entradas al día a casi cero mientras el estado de pools, datasets y snapshots sigue actualizándose.

### v2.9.19

- **Fix del instalador: error de variable no definida resuelto**: el script de instalación ya no falla con el error "existing_webhook: unbound variable" durante la configuración. El WEBHOOK_SECRET ahora se lee correctamente del archivo env.
Expand Down
4 changes: 4 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -359,6 +359,10 @@ Go dependencies (kept to 2 on purpose):

## Changelog

### v2.9.20

- **Reduce sudo/journald log volume from periodic collectors (#124)**: the EasyZFS service generated tens of thousands of sudo log entries per day because every `zpool`/`zfs` command run by the collectors was logged by `sudo` as an allowed command. Add `Defaults:easyzfs !pam_session, !log_allowed` to the sudoers file so allowed commands are no longer logged; failed attempts still are. Complement this with TTL caches in `ZpoolCollector` for stable pool properties (`autotrim`, `checkpoint`, `compressratio`) and for `zpool status -t`, plus a new `EASYZFS_ZPOOL_INTERVAL` environment variable (default 60 s) to control the collector tick interval. Validated on real Proxmox nodes: `journalctl -g 'easyzfs :'` drops from thousands of entries per day to near zero while pool status, datasets and snapshots keep updating.

### v2.9.19
- **Installer fix: unbound variable error resolved**: the installer script no longer fails with an "existing_webhook: unbound variable" error during configuration. The WEBHOOK_SECRET is now correctly read from the env file.

Expand Down
14 changes: 10 additions & 4 deletions deploy/easyzfs.sudoers
Original file line number Diff line number Diff line change
Expand Up @@ -2,17 +2,23 @@
# El servicio corre como usuario 'easyzfs' (sin root) y necesita elevar solo
# estos binarios (executil antepone `sudo -n` cuando euid != 0).
#
# Argumentos restringidos al uso real del código:
# Argumentos restringidos al uso real del codigo:
# crontab → solo `-l` (lectura del crontab de root en la vista Tareas);
# sin args, `crontab -e` permitiría ejecutar código como root.
# sin args, `crontab -e` permitiria ejecutar codigo como root.
# hdparm → solo `-y /dev/*` (standby de disco en acciones).
# udisksctl → solo `power-off -b /dev/*` (apagado de disco; fallback de hdparm).
# dd → solo lectura de disco a /dev/null (parpadeo del LED de la
# bahía por actividad I/O; el count fijo evita lecturas largas).
# bahia por actividad I/O; el count fijo evita lecturas largas).
# zpool/zfs/smartctl/lsblk necesitan libertad de argumentos (son la API del
# almacenamiento); el confinamiento extra vive en easyzfs-sysd y la unit.
#
# Instalación:
# Los comandos automatizados del colector generan miles de entradas de sudo al
# dia. EasyZFS ya guarda su propia auditoria en la BD (tabla audit), asi que
# suprimimos aqui el log repetitivo de PAM y el de comandos permitidos para
# no saturar journald. Los intentos fallidos siguen logueandose.
Defaults:easyzfs !pam_session, !log_allowed
#
# Instalacion:
# install -m 0440 -o root -g root deploy/easyzfs.sudoers /etc/sudoers.d/easyzfs
# visudo -cf /etc/sudoers.d/easyzfs # validar sintaxis antes de recargar
easyzfs ALL=(root) NOPASSWD: /usr/sbin/zpool, /usr/sbin/zfs, /usr/sbin/smartctl, /usr/bin/lsblk, /usr/bin/crontab -l, /usr/sbin/hdparm -y /dev/*, /usr/bin/udisksctl power-off -b /dev/*, /usr/bin/dd if=/dev/* of=/dev/null bs=1M count=2048, /usr/local/libexec/easyzfs-sysd
2 changes: 1 addition & 1 deletion deploy/install.sh
Original file line number Diff line number Diff line change
Expand Up @@ -966,7 +966,7 @@ DEMO=1"
info "MODO DEMO activado (DEMO=1): datos de muestra; las mutaciones responden 403 demo_mode."
info "Para pasar a producción: quita DEMO=1 de ${ENV_FILE} y reinicia el servicio."
else
info "Opcionales que puedes añadir: COOKIE_SECURE=1 (tras proxy TLS), RETENTION_DAYS=30, DEMO=1, MOCK=1."
info "Opcionales que puedes añadir: COOKIE_SECURE=1 (tras proxy TLS), RETENTION_DAYS=30, EASYZFS_ZPOOL_INTERVAL=60 (segundos), DEMO=1, MOCK=1."
fi
}

Expand Down
2 changes: 1 addition & 1 deletion internal/collectors/collectors.go
Original file line number Diff line number Diff line change
Expand Up @@ -63,7 +63,7 @@ func Build(cfg *config.Config, d *sql.DB, h *hub.Hub, al *alerts.Alerter) (*Prov
m := NewMock(h, al)
return &Providers{Pools: m, Disks: m, SysTimers: m, Perf: m, Caps: m}, []Collector{m, mant}
}
zc := NewZpoolCollector(d, h, al)
zc := NewZpoolCollector(d, h, al, cfg.ZpoolInterval)
sc := NewSensorsCollector(h)
smc := NewSmartCollector(d, h, al, sc)
ssc := NewSchedSysCollector()
Expand Down
89 changes: 61 additions & 28 deletions internal/collectors/zpool.go
Original file line number Diff line number Diff line change
Expand Up @@ -23,11 +23,13 @@ import (
)

const (
zpoolInterval = 30 * time.Second
zpoolMaxBackoff = 5 * time.Minute
seriesInterval = 10 * time.Minute // persistir series con esta cadencia mínima
historyTTL = 10 * time.Minute // re-leer 'zpool history' como máximo con esta cadencia
historyTimeout = 90 * time.Second // historiales grandes (bigtank: ~20 s / 275 MB)
zpoolIntervalDef = 60 * time.Second
zpoolMaxBackoff = 5 * time.Minute
seriesInterval = 10 * time.Minute // persistir series con esta cadencia minima
historyTTL = 10 * time.Minute // re-leer 'zpool history' como maximo con esta cadencia
historyTimeout = 90 * time.Second // historiales grandes (bigtank: ~20 s / 275 MB)
propTTL = 5 * time.Minute // propiedades estables: autotrim, checkpoint, compressratio
trimTTL = 2 * time.Minute // estado TRIM no cambia tan rapido; reduce llamadas -t
)

// ZpoolCollector — caché de pools, datasets y snapshots.
Expand All @@ -49,24 +51,36 @@ type ZpoolCollector struct {
prevPct map[string]int
lastSeries map[string]time.Time

// refreshCh despierta el bucle Run tras una mutación (autotrim, trim…):
// sin él la UI vería el valor antiguo hasta el próximo tick de 30 s.
// Buffer 1 = debounce: una ráfaga de mutaciones produce UNA recolecta.
// Intervalo entre recolectas periodicas (configurable; #124).
interval time.Duration

// Cache de propiedades estables y de trim para no repetir comandos en
// cada tick del colector.
lastPropsAt map[string]time.Time

// refreshCh despierta el bucle Run tras una mutacion (autotrim, trim...):
// sin el la UI veria el valor antiguo hasta el proximo tick.
// Buffer 1 = debounce: una rafaga de mutaciones produce UNA recolecta.
refreshCh chan struct{}
}

// NewZpoolCollector crea el colector.
func NewZpoolCollector(d *sql.DB, h *hub.Hub, al *alerts.Alerter) *ZpoolCollector {
// NewZpoolCollector crea el colector. interval=0 usa el default de 60 s.
func NewZpoolCollector(d *sql.DB, h *hub.Hub, al *alerts.Alerter, interval time.Duration) *ZpoolCollector {
if interval <= 0 {
interval = zpoolIntervalDef
}
return &ZpoolCollector{
db: d,
h: h,
al: al,
prevStatus: map[string]string{},
prevPct: map[string]int{},
lastSeries: map[string]time.Time{},
history: map[string][]model.HistoryEntry{},
historyAt: map[string]time.Time{},
refreshCh: make(chan struct{}, 1),
db: d,
h: h,
al: al,
interval: interval,
prevStatus: map[string]string{},
prevPct: map[string]int{},
lastSeries: map[string]time.Time{},
lastPropsAt: map[string]time.Time{},
history: map[string][]model.HistoryEntry{},
historyAt: map[string]time.Time{},
refreshCh: make(chan struct{}, 1),
}
}

Expand All @@ -86,7 +100,7 @@ func (c *ZpoolCollector) RefreshSoon() {

// Run — bucle con ticker, backoff tras 3 fallos seguidos (patrón del skill).
func (c *ZpoolCollector) Run(ctx context.Context) {
interval := zpoolInterval
interval := c.interval
t := time.NewTicker(interval)
defer t.Stop()
if err := c.collectOnce(ctx); err != nil {
Expand Down Expand Up @@ -118,9 +132,9 @@ func (c *ZpoolCollector) Run(ctx context.Context) {
c.stale = true
interval = min(2*interval, zpoolMaxBackoff)
t.Reset(interval)
} else if interval != zpoolInterval {
} else if interval != c.interval {
c.stale = false
interval = zpoolInterval
interval = c.interval
t.Reset(interval)
}
}
Expand Down Expand Up @@ -174,11 +188,12 @@ func (c *ZpoolCollector) collectOnce(ctx context.Context) error {
if err != nil {
return err
}
now := time.Now()
for i := range pools {
c.fillStatus(ctx, &pools[i]) // tolerante: degrada, no falla la pasada
c.fillTrim(ctx, &pools[i]) // progreso de TRIM (zpool status -t)
c.fillCompressRatio(ctx, &pools[i])
c.fillPoolProps(ctx, &pools[i])
c.fillTrim(ctx, &pools[i], now)
c.fillCompressRatio(ctx, &pools[i], now)
c.fillPoolProps(ctx, &pools[i], now)
c.resolveVdevPaths(ctx, &pools[i])
}
history := map[string][]model.HistoryEntry{}
Expand Down Expand Up @@ -265,7 +280,12 @@ func (c *ZpoolCollector) listPools(ctx context.Context) ([]model.Pool, error) {
// fillPoolProps — propiedades del pool autotrim y checkpoint
// ('zpool get -Hp -o property,value autotrim,checkpoint <pool>').
// checkpoint vale "-" cuando no hay checkpoint activo.
func (c *ZpoolCollector) fillPoolProps(ctx context.Context, p *model.Pool) {
// Estas propiedades cambian muy poco: se cachean con TTL para reducir sudo.
func (c *ZpoolCollector) fillPoolProps(ctx context.Context, p *model.Pool, now time.Time) {
key := "props:" + p.Name
if time.Since(c.lastPropsAt[key]) < propTTL {
return
}
out, err := executil.Run(ctx, 5*time.Second, "zpool", "get", "-Hp",
"-o", "property,value", "autotrim,checkpoint", p.Name)
if err != nil {
Expand All @@ -283,6 +303,7 @@ func (c *ZpoolCollector) fillPoolProps(ctx context.Context, p *model.Pool) {
p.Checkpoint = f[1] != "-" && f[1] != ""
}
}
c.lastPropsAt[key] = now
}

// fetchHistory — 'zpool history -i <pool>' parseado EN STREAMING (nil si
Expand Down Expand Up @@ -322,7 +343,12 @@ func (c *ZpoolCollector) History(name string) []model.HistoryEntry {
}

// fillCompressRatio — compressratio del dataset raíz del pool como ratio del pool.
func (c *ZpoolCollector) fillCompressRatio(ctx context.Context, p *model.Pool) {
// Propiedad estable: se cachea con TTL para reducir llamadas a zfs get.
func (c *ZpoolCollector) fillCompressRatio(ctx context.Context, p *model.Pool, now time.Time) {
key := "compress:" + p.Name
if time.Since(c.lastPropsAt[key]) < propTTL {
return
}
out, err := executil.Run(ctx, 5*time.Second, "zfs", "get", "-Hp", "-o", "value",
"compressratio", p.Name)
if err != nil {
Expand All @@ -332,6 +358,7 @@ func (c *ZpoolCollector) fillCompressRatio(ctx context.Context, p *model.Pool) {
if n, err := strconv.ParseFloat(v, 64); err == nil {
p.CompRatio = n
}
c.lastPropsAt[key] = now
}

// --- zpool status: JSON (OpenZFS ≥2.2) con fallback a texto ---
Expand Down Expand Up @@ -667,12 +694,18 @@ func (c *ZpoolCollector) parseStatusText(out string, p *model.Pool) {
// fillTrim — progreso de TRIM ('zpool status -t <pool>'; la salida normal no
// lo muestra). Solo rellena Scrub si el pool no tiene scrub/resilver en curso
// (el scan de datos manda sobre el trim en la representación unificada).
func (c *ZpoolCollector) fillTrim(ctx context.Context, p *model.Pool) {
// Se ejecuta con TTL: reduce llamadas sudo cuando no hay trim activo.
func (c *ZpoolCollector) fillTrim(ctx context.Context, p *model.Pool, now time.Time) {
key := "trim:" + p.Name
if time.Since(c.lastPropsAt[key]) < trimTTL {
return
}
out, err := executil.Run(ctx, 15*time.Second, "zpool", "status", "-t", p.Name)
if err != nil {
return
}
c.parseTrimStatus(string(out), p)
c.lastPropsAt[key] = now
}

// parseTrimStatus — líneas 'scan:' de 'zpool status -t':
Expand Down
Loading
Loading