SLG 游戏的 Logic 服务每秒要处理成千上万次玩家数据查询——登录、改名、升级、战斗,每个操作都要读玩家实体。如果每次都查 Redis,延迟从 0.1ms 涨到 2ms;如果穿透到数据库,直接飙到 20ms。slg-go 的 PlayerCache 用三层防线把 99% 的请求挡在了内存里。

三层防线总览

层级 存储 延迟 TTL
L1 本地 sync.Map ~0.001ms 30s
L2 Redis go-redis ~0.5ms 30s
负缓存 sync.Map ~0.001ms 500ms

PlayerCache 结构体

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
// internal/logic/cache/player_cache.go
type PlayerCache struct {
guard *HotKeyGuard // 热点保护
local sync.Map // L1: map[playerID]*localEntry
redis redis.UniversalClient // L2: Redis
ttl time.Duration // 正缓存 TTL(默认 30s)
negTTL time.Duration // 负缓存 TTL(默认 500ms)
getTTL time.Duration // Redis 读超时(默认 300ms)
sf singleflight.Group // 防穿透
}

type localEntry struct {
player *model.PlayerEntity
expiresAt time.Time
negative bool // 负缓存标记
}

sync.Map 而非 map + RWMutex,因为玩家 ID 分布极度分散(100 万+ ID),sync.Map 的读路径无锁,适合读多写少场景。

singleflight:防穿透的核心

这是最关键的一层。当 100 个 goroutine 同时请求同一个不存在的玩家时,没有 singleflight 会打 100 次 Redis:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
func (c *PlayerCache) Get(ctx context.Context, playerID int64) (*model.PlayerEntity, error) {
// L1: 本地缓存
if p, hit, err := c.getLocal(playerID, true); hit {
return p, err
}

// singleflight 合并并发请求
key := cacheKey(playerID)
value, err, _ := c.sf.Do(key, func() (any, error) {
// 再查一次本地(可能被其他 goroutine 写入了)
if p, hit, localErr := c.getLocal(playerID, false); hit {
return p, localErr
}
// L2: Redis
redisCtx, cancel := c.newRedisGetContext(ctx)
defer cancel()
p, redisErr := c.getRedis(redisCtx, playerID)
if redisErr == nil {
c.setLocal(playerID, p)
return clonePlayer(p), nil
}
// 负缓存
c.setLocalNegative(playerID)
return nil, ErrCacheMiss
})
// ...
}

c.sf.Do(key, fn) 保证同一个 key 只有一个 goroutine 执行 fn,其他 goroutine 等待并共享结果。100 次并发请求 → 1 次 Redis 查询。

独立短超时上下文

singleflight 有个坑:如果首个请求的 ctx 被取消了,后续等待的 goroutine 也会拿到错误。newRedisGetContext 创建独立上下文:

1
2
3
4
5
6
7
8
func (c *PlayerCache) newRedisGetContext(ctx context.Context) (context.Context, context.CancelFunc) {
timeout := c.getTTL // 默认 300ms
base := context.Background()
if ctx != nil {
base = context.WithoutCancel(ctx) // 去掉父 ctx 的取消信号
}
return context.WithTimeout(base, timeout)
}

context.WithoutCancel 是 Go 1.21 新增的 API,创建一个不会被父 ctx cancel 的子 ctx。这保证了即使调用方 ctx 超时,Redis 查询仍能在自己的 300ms 窗口内独立完成。

负缓存:防穿透的第二道保险

当 Redis 也查不到时,不是直接返回,而是写入一个 500ms 的负缓存:

1
2
3
4
5
6
7
func (c *PlayerCache) setLocalNegative(playerID int64) {
if c.negTTL <= 0 { return }
c.local.Store(playerID, &localEntry{
negative: true,
expiresAt: time.Now().Add(c.negTTL), // 500ms
})
}

为什么 500ms 而不是 30s?因为玩家可能刚注册,数据还没写入 Redis。500ms 后自动失效,下次请求会重新查 Redis。如果设太长,新玩家登录会一直报”玩家不存在”。

热点保护:HotKeyGuard

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
// internal/logic/cache/hotkey_guard.go
type HotKeyGuard struct {
mu sync.Mutex
window time.Duration // 滑动窗口(默认 1s)
thresholdQPS int // 热点阈值(默认 1000 QPS)
counters map[string]*counter
}

func (h *HotKeyGuard) Hit(key string) bool {
now := time.Now()
h.mu.Lock()
defer h.mu.Unlock()
c, ok := h.counters[key]
if !ok || now.Sub(c.at) > h.window {
h.counters[key] = &counter{at: now, hits: 1}
return false
}
c.hits++
qps := int(float64(c.hits) / h.window.Seconds())
return qps >= h.thresholdQPS
}

当某个 playerID 的 miss QPS 超过 1000 时,Hit() 返回 true,触发热点保护逻辑(比如直接拒绝查询、告警运维)。这防止了恶意用户用不存在的 playerID 刷爆缓存。

clone 副本:防并发修改

每次 Get() 返回的都是 clonePlayer(p)——值拷贝一份:

1
2
3
4
5
func clonePlayer(p *model.PlayerEntity) *model.PlayerEntity {
if p == nil { return nil }
cp := *p // 浅拷贝
return &cp
}

调用方拿到的是独立副本,修改不会影响缓存里的原始数据。这避免了”读到缓存里的指针 → 修改字段 → 缓存被污染”的经典 bug。

配置项:环境变量驱动

所有参数都通过环境变量配置,无需改代码:

环境变量 默认值 说明
SLG_REDIS_ADDR 空(禁用 Redis) Redis 地址
SLG_CACHE_TTL_SEC 30 正缓存 TTL(秒)
SLG_CACHE_NEG_TTL_MS 500 负缓存 TTL(毫秒)
SLG_CACHE_REDIS_GET_TIMEOUT_MS 300 Redis 读超时(毫秒)

Redis 不可用时自动降级为纯本地缓存(sync.Map),不会阻塞服务。

性能数据

场景 延迟 说明
L1 本地命中 ~0.001ms sync.Map.Load() 无锁读
L2 Redis 命中 ~0.5ms 网络往返
singleflight 合并 ~0.5ms 100 并发 → 1 次 Redis
负缓存命中 ~0.001ms 500ms 内不重复查 Redis
穿透到 DB ~20ms 仅首次 miss

99% 的请求走 L1 本地缓存,0.9% 走 L2 Redis,0.1% 穿透到数据库。

总结

三层防线的设计哲学:

防线 解决的问题 核心技术
L1 本地缓存 高频读取 sync.Map 无锁读
singleflight 缓存穿透 合并并发请求
负缓存 空值穿透 500ms 短 TTL
热点保护 恶意刷缓存 滑动窗口 QPS 检测
clone 副本 并发修改污染 值拷贝隔离

缓存不是”加个 Redis 就行”的事。穿透、击穿、雪崩、并发修改,每个坑都要提前填好。

你们项目里玩家缓存是怎么做的?遇到过缓存穿透打爆数据库的情况吗?