SLG(策略模拟游戏)后端的核心挑战是百万在线连接 。单机扛不住,必须分层。slg-go 采用 Gate→Logic→Battle 三层架构,每层独立扩缩容。
三层职责
graph TD
C["客户端 100万连接"] --> G["Gate × 4-8"]
G -->|"HTTP/gRPC"| L["Logic × 40"]
G -->|"gRPC"| B["Battle × 15"]
L -->|"Kafka"| K["消息总线"]
B --> K
K --> L
层
职责
特点
扩容单位
Gate
TCP/WS 连接管理、协议解析、鉴权
无状态,CPU 低
每台 10-25 万连接
Logic
游戏逻辑(建筑、科技、资源)
有状态分片
每台 2.5 万玩家
Battle
战斗模拟(回合制、实时)
无状态,CPU 密集
按并发战斗数
核心接口抽象 层与层之间通过接口通信,不直接依赖具体实现:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 type Session interface { ID() int64 Send(route uint16 , data []byte ) error Close() error RemoteAddr() net.Addr } type ServiceLocator interface { Locate(ctx context.Context, playerID int64 ) (ServiceActor, error ) } type ServiceActor interface { Tell(ctx context.Context, data []byte ) error Ask(ctx context.Context, data []byte ) ([]byte , error ) }
Session 抽象客户端连接,ServiceLocator 做跨节点玩家定位,ServiceActor 提供 Tell/Ask 两种语义。
Gate 层:连接管理 Gate 层的核心是连接池 + 消息转发 。每个客户端连接由一个 goroutine 处理读,一个 goroutine 处理写:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 func (c *Conn) readPump() { defer c.Close() for { route, payload, err := c.readMessage() if err != nil { return } c.proxy.Forward(c.ctx, c.playerID, c.stateID, route, payload) } } func (c *Conn) writePump() { defer c.Close() for msg := range c.sendCh { if err := c.writeMessage(msg); err != nil { return } } }
Gate 不处理游戏逻辑,只做协议解析和消息路由。路由表根据 route 号决定消息发往 Logic 还是 Battle。
Logic 层:有状态分片 玩家数据按 playerID % shardCount 分片到不同 Logic 实例。同一玩家的所有请求都路由到同一个 Logic 节点,保证状态一致性:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 type Shard struct { id int players map [int64 ]*Player mu sync.RWMutex } func (s *Shard) Handle(ctx context.Context, playerID int64 , route uint16 , data []byte ) error { s.mu.RLock() player, ok := s.players[playerID] s.mu.RUnlock() if !ok { player = s.loadPlayer(ctx, playerID) s.mu.Lock() s.players[playerID] = player s.mu.Unlock() } return player.HandleMessage(ctx, route, data) }
分片数固定(通常 64 或 128),通过 Nacos 服务发现获取 Logic 节点列表,playerID % shardCount 决定路由到哪个节点。
Battle 层:无状态模拟 Battle 层不持有玩家状态,所有数据从请求中携带或从 Redis 加载。战斗结束后结果写回 Redis:
1 2 3 4 5 6 7 8 9 10 11 12 13 func (s *Service) SimulateBattle(ctx context.Context, req *BattleRequest) (*BattleResult, error ) { armyA := s.loadArmy(ctx, req.PlayerA) armyB := s.loadArmy(ctx, req.PlayerB) result := s.engine.Calculate(armyA, armyB) s.saveResult(ctx, result) return result, nil }
无状态意味着 Battle 层可以随意扩缩容。战斗高峰时加实例,空闲时减实例。
消息总线:Kafka Logic 和 Battle 之间的异步通信走 Kafka:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 func (s *LogicService) publishEvent(topic string , event []byte ) { s.kafka.Produce(&kafka.Message{ Topic: topic, Value: event, }) } func (s *BattleService) consumeBattles() { for msg := range s.kafka.Consume("battle.requests" ) { req := parseBattleRequest(msg.Value) result, _ := s.SimulateBattle(context.Background(), req) s.kafka.Produce("battle.results" , result.Marshal()) } }
Kafka 解耦了 Logic 和 Battle,Logic 不需要等 Battle 完成,Battle 完成后通过 Kafka 回调结果。
百万在线的连接管理 Gate 层的连接管理是百万在线的关键。每个 TCP 连接占约 4KB 内存(goroutine 栈 + 缓冲区),100 万连接约 4GB。单台 Gate 服务器通常承载 10-25 万连接,4-8 台 Gate 即可覆盖百万在线。
三层架构的本质是关注点分离 :Gate 管连接,Logic 管状态,Battle 管计算。每层可以独立扩缩容,互不影响。