Commit c1ec05a3 authored by David Yang's avatar David Yang

feat(audit): add AI specialist analysis and shared audit skill

- Add AI-driven audit analysis engine with batched semantic selection
- Introduce shared audit skill (public markdown file) with business
  filter rules, skill upload in prompt config dialog, and rule matching
- Add risk point "exclude and remember" action writing back to skill
- Support configurable Word report style (fonts, spacing, margins)
- Include AI analysis summary and skill rules in Markdown/Excel/Word exports
- Persist skill object via DataService (ai_skill subject)
parent 10603edb
......@@ -15,3 +15,5 @@
.idea/*
.git/*
.DS_Store
.pnpm-store/
This diff is collapsed.
# 审计专项公共 Skill
本文件由历史项目中的以下 3 个 Skill 合并而成:
- `audit_checklist_format.md`
- `retrieval_rules.md`
- `keyword_expansion_rules.md`
## 审计清单格式规则
### 总原则
- 风险点、法规依据、风险点来源、风险等级必须来自本地 Excel 知识库,不得编造。
- 不得改写知识库中的“具体风险点描述”“法规依据”“风险点来源”“风险等级”原文。
- 可以对展示结构、层级标题、字段顺序进行优化,但不得改变事实内容。
- 每条风险点必须展示与用户输入审计项目的匹配度。
### 输出层级
审计内容必须按以下层级组织:
1. 业务板块
2. 一级项目
3. 二级项目
4. 风险点
相同业务板块必须合并在同一标题下;同一业务板块下相同一级项目必须合并;同一一级项目下相同二级项目必须合并。
### 审计清单要素
每次生成审计检查清单时,应包含以下要素:
- 审计项目
- 审计对象
- 审计内容
- 审计方法
- 审计人员
- 审计时间
- 注意事项
### 风险点字段
每个风险点必须包含:
- 匹配度
- 具体风险点描述
- 法规依据
- 风险点来源
- 风险等级
### 展示要求
#### 主界面摘要展示
- 审计项目应该展示业务人员输入的审计项目。
- 审计对象根据审计项目和审计内容进行简单提炼,参考公募基金公司组织架构列示部门。
- 对话界面中,审计内容按“业务板块 → 一级项目 → 二级项目”层级展示。
- 每个二级项目旁仅展示风险点个数(如“共 N 个风险点”),不直接展开所有风险点详情。
- 业务板块、一级项目可使用折叠/展开结构展示不同颗粒度。
- 审计方法按照可能涉及的审计技术手段进行列示。
#### 查看详情
- 每个二级项目提供“查看详情”展开面板,用户点击后可查看该二级项目下的全部风险点完整信息。
- 详情面板中按顺序展示每个风险点的:匹配度、具体风险点描述、法规依据、风险点来源、风险等级。
#### 匹配度排序与标注
- 优先展示匹配度高的风险点。
- 匹配度较低但被保留的风险点也应显示分数,便于业务人员判断是否采用。
- 如果知识库中没有足够相关内容,应明确提示知识库覆盖不足,而不是强行编造。
### Word 导出规则
- 导出 Word 文档前,系统要求用户输入最低匹配度阈值,界面显示为“请输入审计方案中想列示风险点的最低分值”。
- 用户可先通过“查看详情”浏览各风险点分值,确定合适的阈值后再输入。
- 仅匹配度 ≥ 用户设定阈值的风险点纳入 Word 文档。
- Word 文档中审计内容部分按“业务板块 → 一级项目 → 二级项目 → 风险点”完整展示(含通过阈值筛选的全部字段)。
## 检索与评分规则
### 总原则
- 每次检索必须对知识库中的**所有**文档进行评分,不得仅对 top-k 子集评分后截断。
- 匹配度低于阈值的检查点不在主结果中展示,但仍保留在会话状态中供后续调整指令使用。
### 匹配度阈值
- `MATCH_THRESHOLD: 80`
- 匹配度 ≥ 80 的检查点纳入最终输出结果。
- 匹配度 < 80 的检查点保留在 `all_scored_results` 中,供对话修正指令(保留/去掉/增加)操作。
### 评分公式
综合分 = 向量余弦相似度 + 来源约束加分 + 类别命中加分 + 关键词命中加分
#### 来源约束
- 当用户输入包含“来自 X”“来源为 X”“与 X 有关”等表达时,X 应作为风险点来源约束。
- 有明确来源约束时,应先限定“风险点来源”包含该来源的检查点,再进行主题相关性评分。
- 来源约束命中:`+0.36`
- 来源约束命中且同时命中有效主题关键词时,综合分至少提升至 `1.08`,确保直接匹配行不因向量分偏低被漏召回。
#### 类别加分
- 业务板块命中:`+0.18`
- 一级项目命中:`+0.12`
- 二级项目命中:`+0.12`
#### 关键词加分
- 二级项目包含关键词:`+0.12`
- 一级项目或业务板块包含关键词:`+0.09`
- 风险点来源包含关键词:`+0.16`
- 文本内容包含关键词:`+0.06`
#### 匹配度归一化
- 综合分 → 匹配度 = `min(100, round(综合分 / 1.3 * 100))`
- ≥85:高
- ≥70:较高
- ≥55:中
- <55:低
### 检索策略
1. 对全部文档计算向量余弦相似度(矩阵乘法,毫秒级)。
2. 如果用户明确指定来源,先按“风险点来源”字段进行硬过滤。
3. 对过滤后的文档计算来源、类别和关键词加分。
4. 合并得到综合分,归一化为 0-100 匹配度。
5. 按匹配度降序排列,筛选 ≥ 80 的文档作为输出结果。
6. 全部评分文档(含低于阈值的)存入 `all_scored_results` 供调整指令使用。
### 展示要求
- 主结果摘要中标注匹配度 ≥ 80 的条目数量与总检索条数,例如:“检索到 1450 条,匹配度 ≥ 80 分的有 47 条”。
- 审计内容摘要中,每个二级项目仅显示风险点个数,不展开详情。
- 用户可通过“查看详情”展开面板查看每个二级项目下的完整风险点信息(含匹配度、描述、法规依据、来源、等级)。
- 调整指令(如“去掉信息技术相关的”)基于 `all_scored_results` 操作,而非仅基于上次输出的子集。
### Word 导出阈值
- 导出 Word 文档时,系统要求用户设定最低匹配度阈值(默认为 `MATCH_THRESHOLD`)。
- 仅匹配度 ≥ 用户设定阈值的风险点被纳入 Word 文档输出。
- 用户应在导出前通过“查看详情”浏览各风险点分值,确定合适的阈值。
## LLM 关键词扩展规则
### 目标
根据用户输入的审计项目或问题,生成用于本地知识库检索和重排的扩展关键词。
### 关键词生成原则
- 关键词必须围绕用户输入的审计主题。
- 优先生成基金行业、审计、合规、风险管理语境下的专业词。
- 应包含同义词、近义词、监管常用表述、风险点常见表述。
- 关键词用于检索和排序,不得直接作为事实输出。
### 数量与长度
- 生成 10-20 个关键词。
- 每个关键词建议 2-8 个字。
- 用逗号分隔。
- 不要编号,不要解释。
### 示例
#### 反洗钱专项审计
可扩展为:
反洗钱, 洗钱风险, 可疑交易, 大额交易, 客户身份识别, 身份尽职调查, 客户风险等级, 黑名单, 制裁名单, 交易监控, 反洗钱培训, 监管报送
#### 员工行为管理专项审计
可扩展为:
员工行为, 从业人员, 投资行为, 通讯工具, 信息保密, 廉洁从业, 利益输送, 人员资质, 任免管理, 薪酬管理, 问责管理, 员工培训
#### 采购管理专项审计
可扩展为:
采购, 供应商, 招标, 询价, 合同, 验收, 付款, 服务机构, 外包, 第三方, 廉洁从业, 服务协议
### 禁止事项
- 不要生成与用户审计主题无关的宽泛词。
- 不要输出完整句子。
- 不要输出结论或法规解释。
This diff is collapsed.
export const AUDIT_KNOWLEDGE_FILE_SUBJECT = 'audit_knowledge_file';
export const AUDIT_SKILL_SUBJECT = 'ai_skill';
export const AUDIT_MATCH_THRESHOLD = 80;
export const AUDIT_PROMPT_RECORD_LIMIT = 80;
......
This diff is collapsed.
This diff is collapsed.
This diff is collapsed.
export const DEFAULT_WORD_STYLE = {
fontFamily: '宋体',
bodyFontSizePt: 10.5,
metaFontSizePt: 9,
heading1FontFamily: '黑体',
heading1FontSizePt: 14,
heading2FontFamily: '黑体',
heading2FontSizePt: 12,
heading3FontFamily: '黑体',
heading3FontSizePt: 11,
lineSpacing: 1.15,
paragraphAfterPt: 6,
firstLineIndentChars: 2,
pageMarginsMm: {
top: 25.4,
right: 25.4,
bottom: 25.4,
left: 25.4,
},
};
const FONT_NAMES = ['宋体', '黑体', '仿宋', '楷体', '微软雅黑', '等线'];
const SIZE_ALIASES = {
初号: 42,
小初: 36,
一号: 26,
小一: 24,
二号: 22,
小二: 18,
三号: 16,
小三: 15,
四号: 14,
小四: 12,
五号: 10.5,
小五: 9,
};
function clampNumber(value, min, max, fallback) {
const number = Number(value);
if (!Number.isFinite(number)) return fallback;
return Math.min(max, Math.max(min, number));
}
function normalizeFont(value, fallback) {
const text = String(value || '').trim();
return FONT_NAMES.find((font) => text.includes(font)) || fallback;
}
export function normalizeWordStyle(style = {}) {
const source = style || {};
const margins = source.pageMarginsMm || source.pageMargins || {};
return {
...DEFAULT_WORD_STYLE,
...source,
fontFamily: normalizeFont(source.fontFamily, DEFAULT_WORD_STYLE.fontFamily),
bodyFontSizePt: clampNumber(
source.bodyFontSizePt,
8,
24,
DEFAULT_WORD_STYLE.bodyFontSizePt,
),
metaFontSizePt: clampNumber(
source.metaFontSizePt,
7,
18,
DEFAULT_WORD_STYLE.metaFontSizePt,
),
heading1FontFamily: normalizeFont(
source.heading1FontFamily,
DEFAULT_WORD_STYLE.heading1FontFamily,
),
heading1FontSizePt: clampNumber(
source.heading1FontSizePt,
10,
30,
DEFAULT_WORD_STYLE.heading1FontSizePt,
),
heading2FontFamily: normalizeFont(
source.heading2FontFamily,
DEFAULT_WORD_STYLE.heading2FontFamily,
),
heading2FontSizePt: clampNumber(
source.heading2FontSizePt,
9,
26,
DEFAULT_WORD_STYLE.heading2FontSizePt,
),
heading3FontFamily: normalizeFont(
source.heading3FontFamily,
DEFAULT_WORD_STYLE.heading3FontFamily,
),
heading3FontSizePt: clampNumber(
source.heading3FontSizePt,
9,
24,
DEFAULT_WORD_STYLE.heading3FontSizePt,
),
lineSpacing: clampNumber(
source.lineSpacing,
1,
3,
DEFAULT_WORD_STYLE.lineSpacing,
),
paragraphAfterPt: clampNumber(
source.paragraphAfterPt,
0,
36,
DEFAULT_WORD_STYLE.paragraphAfterPt,
),
firstLineIndentChars: clampNumber(
source.firstLineIndentChars,
0,
4,
DEFAULT_WORD_STYLE.firstLineIndentChars,
),
pageMarginsMm: {
top: clampNumber(
margins.top,
10,
50,
DEFAULT_WORD_STYLE.pageMarginsMm.top,
),
right: clampNumber(
margins.right,
10,
50,
DEFAULT_WORD_STYLE.pageMarginsMm.right,
),
bottom: clampNumber(
margins.bottom,
10,
50,
DEFAULT_WORD_STYLE.pageMarginsMm.bottom,
),
left: clampNumber(
margins.left,
10,
50,
DEFAULT_WORD_STYLE.pageMarginsMm.left,
),
},
};
}
function parseSize(value, fallback) {
const text = String(value || '').trim();
const alias = Object.keys(SIZE_ALIASES).find((key) => text.includes(key));
if (alias) return SIZE_ALIASES[alias];
const match = text.match(/(\d+(?:\.\d+)?)\s*(?:pt|磅|号)?/i);
return match ? Number(match[1]) : fallback;
}
function parseLineSpacing(value, fallback) {
const text = String(value || '').trim();
if (text.includes('单倍')) return 1;
if (text.includes('双倍') || text.includes('两倍')) return 2;
const match = text.match(/(\d+(?:\.\d+)?)\s*(?:倍|行)?/i);
return match ? Number(match[1]) : fallback;
}
function parsePoint(value, fallback) {
const match = String(value || '').match(/(\d+(?:\.\d+)?)/);
return match ? Number(match[1]) : fallback;
}
export function parseWordStyleMarkdown(
markdown = '',
fallback = DEFAULT_WORD_STYLE,
) {
const style = { ...normalizeWordStyle(fallback) };
const source = String(markdown || '');
const read = (pattern) => source.match(pattern)?.[1]?.trim() || '';
const bodyFont = read(/(?:正文字体|正文用字体)\s*[::]\s*([^\n]+)/i);
const bodySize = read(
/(?:正文(?:字号|大小)|正文字体大小)\s*[::]\s*([^\n]+)/i,
);
const heading1 = read(/(?:一级标题(?:字体|字号|大小)?)\s*[::]\s*([^\n]+)/i);
const heading2 = read(/(?:二级标题(?:字体|字号|大小)?)\s*[::]\s*([^\n]+)/i);
const heading3 = read(/(?:三级标题(?:字体|字号|大小)?)\s*[::]\s*([^\n]+)/i);
const line = read(/(?:行距|行间距)\s*[::]\s*([^\n]+)/i);
const after = read(/(?:段后|段后间距)\s*[::]\s*([^\n]+)/i);
const indent = read(/(?:首行缩进|首行)\s*[::]\s*([^\n]+)/i);
if (bodyFont) style.fontFamily = normalizeFont(bodyFont, style.fontFamily);
if (bodySize)
style.bodyFontSizePt = parseSize(bodySize, style.bodyFontSizePt);
if (heading1) {
style.heading1FontFamily = normalizeFont(
heading1,
style.heading1FontFamily,
);
style.heading1FontSizePt = parseSize(heading1, style.heading1FontSizePt);
}
if (heading2) {
style.heading2FontFamily = normalizeFont(
heading2,
style.heading2FontFamily,
);
style.heading2FontSizePt = parseSize(heading2, style.heading2FontSizePt);
}
if (heading3) {
style.heading3FontFamily = normalizeFont(
heading3,
style.heading3FontFamily,
);
style.heading3FontSizePt = parseSize(heading3, style.heading3FontSizePt);
}
if (line) style.lineSpacing = parseLineSpacing(line, style.lineSpacing);
if (after) style.paragraphAfterPt = parsePoint(after, style.paragraphAfterPt);
if (indent)
style.firstLineIndentChars = parsePoint(indent, style.firstLineIndentChars);
return normalizeWordStyle(style);
}
export function isAuditStyleRequest(query = '') {
return /(字体|字号|大小|行距|行间距|段前|段后|缩进|页边距|排版|格式|样式)/.test(
String(query),
);
}
export function isPersistentStyleRequest(query = '') {
return /(以后|今后|默认|保存为|记住|都这样|统一样式|长期)/.test(
String(query),
);
}
export function parseAuditStyleRequest(
query = '',
currentStyle = DEFAULT_WORD_STYLE,
) {
const text = String(query || '').trim();
const current = normalizeWordStyle(currentStyle);
if (!isAuditStyleRequest(text)) {
return { matched: false, style: current, changes: [], persist: false };
}
const style = { ...current };
const changes = [];
const add = (key, value, label) => {
if (value === undefined || value === null || value === style[key]) return;
style[key] = value;
changes.push(`${label}:${value}`);
};
const fontMatch = text.match(
/(?:字体|字形)(?:改成|设置为|调整为|用|为)?\s*(宋体|黑体|仿宋|楷体|微软雅黑|等线)/,
);
if (fontMatch) add('fontFamily', fontMatch[1], '正文字体');
const bodySizeMatch = text.match(
/正文(?:字体|字号|大小)?(?:改成|设置为|调整为|用|为)?\s*(初号|小初|一号|小一|二号|小二|三号|小三|四号|小四|五号|小五|\d+(?:\.\d+)?\s*(?:pt|磅|号))/,
);
if (bodySizeMatch)
add(
'bodyFontSizePt',
parseSize(bodySizeMatch[1], style.bodyFontSizePt),
'正文字号',
);
const headingMatch = text.match(
/(?:一级标题|标题1)(?:字体|字号|大小)?(?:改成|设置为|调整为|用|为)?\s*(?:(宋体|黑体|仿宋|楷体|微软雅黑|等线)\s*)?(初号|小初|一号|小一|二号|小二|三号|小三|四号|小四|五号|小五|\d+(?:\.\d+)?\s*(?:pt|磅|号))/,
);
if (headingMatch) {
if (headingMatch[1])
add('heading1FontFamily', headingMatch[1], '一级标题字体');
add(
'heading1FontSizePt',
parseSize(headingMatch[2], style.heading1FontSizePt),
'一级标题字号',
);
}
const lineMatch = text.match(
/(?:行距|行间距)(?:改成|设置为|调整为|为)?\s*(单倍|双倍|两倍|\d+(?:\.\d+)?\s*倍?)/,
);
if (lineMatch)
add(
'lineSpacing',
parseLineSpacing(lineMatch[1], style.lineSpacing),
'行距',
);
const afterMatch = text.match(
/(?:段后|段后间距)(?:空|设置为|调整为|为)?\s*(\d+(?:\.\d+)?)\s*(?:磅|pt)?/,
);
if (afterMatch) add('paragraphAfterPt', Number(afterMatch[1]), '段后');
const indentMatch = text.match(
/(?:首行缩进|首行)(?:设置为|调整为|为)?\s*(\d+(?:\.\d+)?)\s*(?:字符|字)?/,
);
if (indentMatch)
add('firstLineIndentChars', Number(indentMatch[1]), '首行缩进');
return {
matched: true,
style: normalizeWordStyle(style),
changes,
persist: isPersistentStyleRequest(text),
};
}
export function wordStyleToMarkdown(style = DEFAULT_WORD_STYLE) {
const normalized = normalizeWordStyle(style);
return [
'## Word 输出样式',
'',
'### 公共默认样式',
`- 正文字体:${normalized.fontFamily}`,
`- 正文大小:${normalized.bodyFontSizePt}pt`,
`- 一级标题:${normalized.heading1FontFamily},${normalized.heading1FontSizePt}pt`,
`- 二级标题:${normalized.heading2FontFamily},${normalized.heading2FontSizePt}pt`,
`- 三级标题:${normalized.heading3FontFamily},${normalized.heading3FontSizePt}pt`,
`- 行距:${normalized.lineSpacing} 倍`,
`- 段后:${normalized.paragraphAfterPt} 磅`,
`- 首行缩进:${normalized.firstLineIndentChars} 字符`,
].join('\n');
}
This diff is collapsed.
......@@ -356,6 +356,31 @@ export async function uploadAuditKnowledgeFile(file) {
return normalizeUploadedFile(payload, file);
}
export async function uploadAuditSkillFile(
content,
fileName = 'audit_public.md',
) {
const file =
content instanceof File
? content
: new File([String(content || '')], fileName, {
type: 'text/markdown;charset=utf-8',
lastModified: Date.now(),
});
const formData = new FormData();
formData.append('file', file, file.name);
const response = await fetch('/api/upload', {
method: 'POST',
body: formData,
credentials: 'include',
});
if (!response.ok) {
throw new Error(`Skill 文件上传失败(${response.status})`);
}
const payload = await response.json();
return normalizeUploadedFile(payload, file);
}
export function getUserInfo() {
const currentUser = localStorage.getItem('current_user');
return currentUser && JSON.parse(currentUser);
......
Markdown is supported
0% or
You are about to add 0 people to the discussion. Proceed with caution.
Finish editing this message first!
Please register or to comment