easy-word 的数据层要解决两件事:单词要同时展示 IPA 音标和简易拼读标注,多设备之间要保持学习进度同步。这两个问题的解法都不复杂,但有些细节值得记录。

双音标系统

模型定义

单词的核心数据结构在 entry/src/main/ets/model/Word.ets:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
export interface Word {
id: string;
word: string;
phoneticIPA: string; // IPA 国际音标,如 "/həlˈoʊ/"
phoneticSimple: string; // 简易拼读标注,如 "hel-lo"
translation: string;
partOfSpeech: string;
exampleEn: string;
exampleEnAlt: string;
exampleCn: string;
audioUrl: string | null;
grade: 1 | 2 | 3 | 4 | 5 | 6;
unit: number;
difficulty: 1 | 2 | 3;
tags: string[];
}

phoneticIPA 和 phoneticSimple 是两套互补的标注:

  • IPA 音标:标准国际音标,面向有音标基础的用户
  • 简易拼读:按音节拆分的自然拼读标注,面向低年级学生

JSON 数据文件中的实际值:

1
2
3
4
{
"phoneticIPA": "/həlˈoʊ/",
"phoneticSimple": "hel-lo"
}

UI 渲染

在 WordFlashCard 组件中,两套音标分开展示:

1
2
3
4
5
6
7
8
9
10
11
12
Row({ space: 10 }) {
Text(this.word.phoneticIPA)
.fontSize(16)
.fontColor(this.isDark() ? this.getColors().textSecondary : '#666666')
Button('🔊')
}
if (this.word.phoneticSimple) {
Text(this.word.phoneticSimple)
.fontSize(18)
.fontColor(this.isDark() ? this.getColors().textSecondary : '#F5A623')
.fontWeight(FontWeight.Medium)
}

IPA 音标用灰色小字显示,简易拼读用橙色高亮。这样设计的考虑是:低年级学生可能还没学过 IPA,简易拼读更直观;而高年级学生可以参考 IPA 做精确发音。

发音切换

发音服务在 entry/src/main/ets/service/PronunciationService.ets 中支持美式和英式两种口音:

1
2
3
4
5
6
7
private currentAccent: 'us' | 'uk' = 'us';

private applySpeechProfile(accent: 'us' | 'uk', speed: number): void {
this.currentAccent = accent;
this.currentSpeechSpeed = this.normalizeSpeed(speed);
this.currentSpeechPitch = accent === 'uk' ? 0.94 : 1.0;
}

英式发音的音调参数略低(0.94),模拟英式英语的语调特征。用户可以在设置页切换口音偏好。

数据库模型

RDB 表结构

easy-word 使用鸿蒙 NEXT 的 relationalStore 关系型数据库,数据库文件名为 wordcard.db。整个应用有 11 张需要分布式同步的表:

表名 用途 主键
child_profiles 学习者档案 id
word_progress 单词学习进度 (child_id, word_id)
check_ins 每日打卡记录 (child_id, date)
badges 徽章获得记录 (child_id, badge_id)
user_settings 用户设置 union_id
pets 虚拟宠物 child_id
pet_accessories 宠物配饰 id
quiz_sessions 测验会话 id
quiz_records 测验记录 id
phonics_game_records 自然拼读游戏记录 id
learning_sessions 学习时长记录 id

核心表:word_progress

这是最复杂的一张表,同时承载学习进度和艾宾浩斯复习调度:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
CREATE TABLE IF NOT EXISTS word_progress (
child_id TEXT NOT NULL,
word_id TEXT NOT NULL,
seen_count INTEGER DEFAULT 0,
correct_count INTEGER DEFAULT 0,
last_seen_at INTEGER,
mastery_level INTEGER DEFAULT 0,
ebb_interval INTEGER DEFAULT 0, -- 艾宾浩斯间隔(分钟)
ebb_ease REAL DEFAULT 2.5, -- 难度系数
ebb_next_at INTEGER DEFAULT 0, -- 下次复习时间戳
ebb_repetitions INTEGER DEFAULT 0, -- 已重复次数
updated_at INTEGER NOT NULL, -- 本地修改时间
synced_at INTEGER DEFAULT 0, -- 上次同步时间
phonics_mastery INTEGER DEFAULT 0, -- 自然拼读掌握度
PRIMARY KEY (child_id, word_id)
)

ebb_ 前缀的字段是艾宾浩斯复习调度算法的状态。updated_at 和 synced_at 是双时间戳同步机制的核心。

DAO 层设计

数据库访问采用 DAO 模式,RdbHelper 是单例门面,聚合了所有 DAO:

1
2
3
4
5
6
7
8
9
10
11
12
13
class RdbHelper {
private static instance: RdbHelper;
private store: relationalStore.RdbStore | null = null;

readonly childProfileDao: ChildProfileDao;
readonly wordProgressDao: WordProgressDao;
readonly checkInDao: CheckInDao;
readonly badgeDao: BadgeDao;
readonly userSettingsDao: UserSettingsDao;
readonly petDao: PetDao;
readonly quizDao: QuizDao;
readonly learningSessionDao: LearningSessionDao;
}

每个 DAO 负责一张表的 CRUD 操作。DAO 之间不互相调用,跨表查询在 Service 层组合。

双时间戳分布式同步

核心思路

每条记录维护两个时间戳:

  • updated_at:本地最后一次修改的时间
  • synced_at:上次成功同步到远端设备的时间

判断数据是否需要同步的条件非常简单:updated_at > synced_at。满足这个条件的记录就是”脏数据”,需要推送到其他设备。

脏数据查询

所有 DAO 的脏数据查询都是同一个模式:

1
2
3
4
5
6
async getDirtyWordProgress(childId?: string): Promise<WordProgress[]> {
const sql = childId
? 'SELECT * FROM word_progress WHERE updated_at > synced_at AND child_id = ?'
: 'SELECT * FROM word_progress WHERE updated_at > synced_at';
// ...
}

Upsert 保留 synced_at

这是整个同步机制中最需要注意的细节。当从远端拉取数据写入本地时,需要用 INSERT OR REPLACE。但 INSERT OR REPLACE 会清空所有字段,包括 synced_at——如果不处理,刚同步过来的记录会被标记为脏数据,下次推送时又推回去,形成死循环。

解决方案是用 COALESCE 保留旧的 synced_at:

1
2
3
INSERT OR REPLACE INTO child_profiles (id, nickname, ..., synced_at)
VALUES (?, ?, ?, ?, ?, ?,
COALESCE((SELECT synced_at FROM child_profiles WHERE id = ?), 0))

COALESCE 的逻辑是:如果本地已经存在这条记录,取旧的 synced_at;如果不存在(新记录),默认为 0(未同步)。这样保证了拉取操作不会意外标记脏数据。

同步后标记

推送成功后,把所有脏记录的 synced_at 更新为当前时间:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
async markWordProgressSynced(
records: Array<WordProgressSyncRecord>,
syncedAt: number
): Promise<void> {
await db.executeSql('BEGIN TRANSACTION');
try {
for (const r of records) {
await db.executeSql(
'UPDATE word_progress SET synced_at = ? WHERE child_id = ? AND word_id = ?',
[syncedAt, r.childId, r.wordId]
);
}
await db.executeSql('COMMIT');
} catch (error) {
await db.executeSql('ROLLBACK');
throw new Error(String(error));
}
}

用事务包裹批量更新,保证原子性。

同步流程

同步服务在 entry/src/main/ets/service/CloudSyncService.ets 中实现,采用 ArkData 分布式 RDB 的推拉模型:

1
2
3
4
5
6
7
8
9
10
11
12
13
async pushDirty(): Promise<SyncSummary> {
return this.syncTables(relationalStore.SyncMode.SYNC_MODE_PUSH, '设备推送');
}

async pullLatest(): Promise<SyncSummary> {
return this.syncTables(relationalStore.SyncMode.SYNC_MODE_PULL, '设备拉取');
}

async fullSync(): Promise<SyncSummary> {
const push = await this.pushDirty();
const pull = await this.pullLatest();
return { ok: push.ok && pull.ok, mode: 'full', ... };
}

逐表同步,每张表独立判断成功或失败,某张表失败不影响其他表:

1
2
3
4
5
6
7
for (const table of SYNC_TABLES) {
const predicates = new relationalStore.RdbPredicates(table);
predicates.inDevices(deviceIds);
const result = await store.sync(mode, predicates);
const tableOk = result.every((item) => item[1] === 0);
if (tableOk) { successCount++; } else { failedCount++; }
}

冲突解决

项目采用 ArkData 内置的 Last-Write-Wins(LWW)策略——如果两台设备同时修改了同一行,updated_at 更大的版本胜出。应用层不做额外的冲突合并。

减少冲突窗口的方式是:App 切回前台时自动触发一次 pull。

1
2
3
4
async syncOnForeground(): Promise<void> {
// App 从后台切到前台时调用
await this.pullLatest();
}

这样用户在设备 A 上修改了学习进度,切到设备 B 时会先拉取最新数据,再开始本地修改,减少了两端同时修改同一行的概率。

设计取舍

这套方案的取舍比较直白。信任 ArkData 的 LWW 冲突解决,不做应用层合并。单词学习进度这种场景,少量数据被覆盖可以接受。所有业务逻辑基于本地 RDB 查询,同步是异步后台操作,断网也能用。COALESCE 保 synced_at 那条 SQL 是最值得记住的技巧。

两台设备间同步延迟 1-3 秒,学习进度基本无感。两台以上还没试过,可能得上版本向量。