cache: fix test flakiness by stopping the chunk cleaner promptly

The background chunk cleaner slept for the whole ChunkCleanInterval
(default 1 minute) before checking its stop channel, and only ran
CleanUpCache via the select default branch. This meant a cache that
had been stopped by StopBackgroundRunners could keep running
CleanUpCache for up to an interval afterwards.

The cache backend tests all share a single on-disk chunk store (the
TestInternalCache remote), so a lingering cleaner from a finished test
could call CleanChunksBySize and os.RemoveAll chunks that a later,
unrelated test had just written. The later test would then read a
chunk back and get an unexpected EOF - eg
TestInternalMaxChunkSizeRespected failing intermittently on CI.

Wait on a timer and the stop channel together so a stop is honoured
immediately and the cleaner can never run again once stopped.
This commit is contained in:
Nick Craig-Wood
2026-07-12 13:26:17 +01:00
parent 9a49790797
commit 63439b4444
+6 -2
View File
@@ -503,15 +503,19 @@ func NewFs(ctx context.Context, name, rootPath string, m configmap.Mapper) (fs.F
}
go func() {
// Wait on the timer and the stop channel together so that a stop
// signalled by StopBackgroundRunners is honoured immediately.
timer := time.NewTimer(time.Duration(f.opt.ChunkCleanInterval))
defer timer.Stop()
for {
time.Sleep(time.Duration(f.opt.ChunkCleanInterval))
select {
case <-f.cleanupChan:
fs.Infof(f, "stopping cleanup")
return
default:
case <-timer.C:
fs.Debugf(f, "starting cleanup")
f.CleanUpCache(false)
timer.Reset(time.Duration(f.opt.ChunkCleanInterval))
}
}
}()