Y/YUXIANG WANG

AI PRODUCT CASE STUDY / 01

PROJECT KNOWLEDGE COPILOT

从历史资料中找答案,
让每个判断有据可查。

面向建筑与设计团队的 AI 知识助手。用自然语言查找、比较历史项目,并从回答中的引用回到原始资料,核验每个判断的依据。

RAG产品决策Evaluation可追溯体验可运行 MVP

应用目前为本地 Demo,公开在线体验与演示视频尚未提供。

我的角色

产品定义 · 架构判断 · 评估与验收

项目时间

2026.05 — 2026.08
后续持续迭代

实现方式

本人主导 · AI 辅助开发

问题来源

建筑与设计背景 · 真实项目资料

一分钟概览 / AT A GLANCE

直接查看证据 ↓

用户问题

建筑资料分散;找到相关项目后,仍需判断回答是否有原文支持。

我的关键选择

把检索与生成分开评估。项目均衡策略未补齐有效证据,因此不设为默认。

已有结果

Phase 9 的 4 道比较题共需覆盖 11 次目标项目:检索找到的目标从 9 次增至 11 次,但找到有效证据的仍只有 6 次。

证据边界

小样本真实资料测试;指标为特定评测范围,不代表企业采用或业务效果。

产品与关键取舍 / PRODUCT IN CONTEXT

相关,不等于有证据。

产品链路示意 · 根据已实现流程整理,非界面截图或在线 Demo。

  1. 01

    限定项目提问

    自然语言搜索,必要时选择项目范围

  2. 02

    检索原文片段

    回答使用 Top 5 证据,与项目卡片分开

  3. 03

    生成带引用回答

    不足时拒答;异常不冒充正常结果

  4. 04

    回到原始 PDF

    从引用核验来源,再判断是否可复用

实验改变了什么

找到目标项目的次数

9/11 → 11/11
用户需要的证据呢?

找到有效证据的次数

6/11 → 6/11

结论:不将项目均衡策略设为默认。分母 11 是 4 道比较题(C1–C4)所需的目标项目次数之和,同一项目在不同题中重复计数,不是 11 道题或 11 个不同项目;有效证据使用确定性模式检查,并非逐句人工认证。

产品推导 / CASE STUDY

从问题到选择,再到验证。

01 / THE PROBLEM

资料不断积累,
知识却很难复用。

新方案需要参考案例,投标需要核对材料与空间策略,新成员需要理解项目历史。答案可能已经存在,却分散在设计说明、技术总结和长篇 PDF 中。查找往往依赖文件名、个人记忆,以及熟悉项目的同事。

CORE INSIGHT

找到一段文字,只完成了一半任务。

用户还需要确认它来自哪里、是否适用于当前判断。产品必须把“发现资料”和“核验依据”连成一条完整路径。

这是基于专业背景和真实资料形成的问题假设,尚无正式访谈样本、企业试点或付费客户记录。“减少查找时间”是产品目标,尚未通过计时对照验证。

02 / THE EXPERIENCE

找到、理解,再核验。

01

确定搜索范围

不知道相关项目时搜索全部资料;知道目标时选择一个项目;需要比较时明确选择 2–4 个项目。

02

阅读受证据约束的回答

系统组织检索片段并附上引用,证据不足时说明原因;请求失败与证据不足采用不同状态。

03

回到原始资料

从行内引用查看来源片段与相邻上下文,进入项目资料页,阅读逻辑文档并在 PDF 阅读器中翻页、缩放与切换分卷。

以上为已实现流程的文字说明。案例未使用真实投标资料截图;公开视觉素材将使用独立合成数据。

03 / PRODUCT DECISIONS

最重要的优化,
来自没有改善的结果。

D21 / 检索策略实验

项目覆盖变好了,为什么仍不采用?

全局 Top-5 可能遗漏跨项目证据。我尝试扩大到 Top-20 候选,再按项目均衡选出 5 个 chunks;Embedding、生成模型与提示词保持一致。

C1–C4 对照指标原方案项目均衡方案
找到目标项目(共需 11 次)9 次11 次
找到有效证据(同一组 11 次需求)6 次6 次
C4 所需证据覆盖1/31/3

4 道比较题合计需要覆盖 11 次目标项目,同一项目在不同题中分别计数。找到目标项目表示检索结果中出现该项目;有效证据则还需通过问题对应的确定性模式检查。

决策:保留默认全局检索。

更多项目出现,没有带来更多问题所需的事实。后续转向让用户明确选择比较对象,再在目标项目内检索;选中项目仍不保证证据充分。

D30 / 信息组织与上下文

浏览十个项目,不必让模型读十份资料。

用户以“项目”理解内容,系统却以 chunk 检索。直接把 Top-5 改成 Top-10,仍可能只有少数几个项目。

Top-5 证据

全局回答继续使用少量具体片段,控制上下文规模。

Top-10 项目

按项目聚合结果,提供最多十张不同项目卡片。

两路检索复用同一个问题向量。项目浏览范围与回答证据量分别设计,也避免为展示卡片重复生成问题向量。

D18 / 信任与失败体验

“没有依据”和“请求失败”,需要不同的下一步。

系统区分 answered、no_evidence、unverifiable 和 error,分别对应正常回答、证据不足、引用无法核验与请求错误。原始评估暴露一次结构化生成异常后,我补充了有限重试,原问题定向复测成功。

没有完整重跑原始 16 题,因此不能把定向复测成功写成失败率降至 0%;unverifiable 的前端状态曾使用模拟响应验证。

D29–D32 / 大文档与阅读体验

底层分卷,不应割裂用户的阅读任务。

真实投标 PDF 暴露存储大小与向量化限流约束。我加入入库工作量预览、可恢复批量 Embedding,并将物理分卷组织成逻辑文档,保留原始页码。搜索框前置、项目选择折叠,再用 PDF.js 完成原文阅读。

三个分卷已能呈现为一份 134 页逻辑文档,第二卷第一页对应原始第 46 页。恢复处理会跳过已有向量,减少重复调用。

RAG 架构与产品取舍
RAG 架构与产品取舍

来自用户作品集 PDF 第 4 页,以两倍尺寸重新渲染,展示资料处理、检索、引用与验证的关系。 · 点击放大

04 / SYSTEM DESIGN

让错误,可以被定位。

将文档处理、检索、生成和引用核验拆开检查。Voyage 负责向量化与语义查找,Claude 负责基于证据组织回答,引用编号由程序校验;编号有效仍不等于语义有依据。

INGESTION / 入库
  1. 上传 PDF / TXT / Markdown
  2. 提取文本 · 预览工作量
  3. 分块 · 保存项目、文档与页码
  4. Voyage 批量向量化 → pgvector

目标分块长度 1,000 字符,重叠 150 字符;PDF 先按页切分,再在页内分块。

QUERY / 查询与核验
  1. 提问并选择项目范围
  2. Voyage 生成一次问题向量
  3. 证据检索 + 项目级聚合检索
  4. Claude 生成引用回答 / 说明证据不足
  5. 引用编号校验 → 来源上下文 → 原文阅读

Next.js · TypeScript · Supabase Postgres / 私有 Storage · pgvector · Voyage AI · Claude API · pdf-parse · PDF.js

PDF 阅读能力与模型理解能力分开:当前主要依赖文本提取,不包含 OCR 或图纸视觉检索。

05 / EVIDENCE & EVALUATION

保留分母,也保留失败。

不同阶段使用不同数据与检查方式,以下分别呈现,避免混成一个“系统准确率”。结果摘自项目评估记录,本次网站整理没有重新调用模型评测。

Phase 8 · 小规模真实资料基线

3 个真实个人建筑项目、6 份文档、19 个 chunks;16 题包括 8 个已知答案、4 个跨项目比较和 4 个未知答案。

检查项结果分母与方法
Retrieval Hit@17/8 · 87.5%8 个已知答案问题
Retrieval Hit@38/8正确证据位于前三名
相关性、依据与引用支持10/10仅成功回答;由作者检查
应拒答案例5/54 个未知 + 1 个证据不完整的比较
原始生成失败1/16保留原始失败,后续仅定向复测
成功请求延迟中位数约 6.7 秒该小规模测试中的观察值

小样本、自评且存在生成失败,不能解释为“系统准确率 100%”或普遍的零幻觉保证。

Phase 10 · 大型真实投标 PDF

293 页2 份原始 PDF · 5 份私有分卷289/289chunks 完成 Embedding

入库提取 37,964 字符。7 次文档 Embedding 请求成功,另有 1 次限流失败。一个比较问题在全局检索中仅覆盖 1/2 目标项目,显式选择后覆盖 2/2;两个未知问题均正确返回 no_evidence。

包含定向复测的 6 次成功回答,引用编号有效,并通过 Claude 辅助语义审查。原始运行有 2 个问题在有限重试后仍生成失败,其中 1 个复测成功、另 1 个仍失败。

REMAINING PRODUCT PROBLEM

成功路径约 14.2 秒,失败复测约 200.6 秒。

前者是六次成功回答的延迟中位数,后者是一次失败问题的定向复测耗时。失败路径的等待体验仍需优先改进。模型辅助审查不等同于独立人工认证。

合成 Demo · 数据准备完成,质量评测待完成

独立环境包含 12 个虚构项目、24 份文档、48 个 chunks,全部完成向量化;24/24 原始文件与生成源一致。先定义事实清单,再确定性生成文档,最后校验与入库。

独立测试集包含 24 题,但尚无整套质量评测完成记录。真实资料基线指标不能移用到合成 Demo。

Phase 11 · 功能回归

真实环境返回 6 张不同项目卡片,合成环境能返回 10 张;完成逻辑文档分组、翻页与分卷切换验证。容量检查复用已有向量,无额外模型调用。

十张卡片属于容量验证,不代表项目召回率 100%;目前没有项目级相关性标注集。

在源码仓库查看产品与评估文档 ↗

06 / REFLECTION

交付一个 MVP,
也明确它的边界。

我主导行业问题定义、MVP 范围、架构判断、评估设计与交付验收。Phase 1–8 主要由 Claude Code 协助实现,后续由 Codex 接手开发协作;产品判断通过决策记录与实验结果保留。

当前局限

  • 没有 OCR 或图纸视觉理解,短标签和上下文不足的片段检索较弱。
  • 有限重试仍会失败;入库需分步触发,分卷元数据主要手动维护。
  • 真实资料实例未配置登录与资源授权,保持本地使用。
  • 尚无企业采用、持续使用或节省工时的数据。

下一步验证

  1. 完成合成 Demo 的独立 24 题评估。
  2. 用合成资料制作截图、视频与 PDF 阅读演示。
  3. 开展文件夹搜索与本产品的任务计时对照。
  4. 公开交互前完善访问控制、写入限制与调用成本控制。

这次实践让我更关注:用户是否拿到了所需证据,以及下一步能否核验和使用它。

← 返回全部作品与我交流 ↗

证据入口 / EVIDENCE

结论可以回到记录。

以下为已有项目记录的归档或节选。本轮核对文件与结论,未重新运行模型评测;原始运行、定向复测和用户验证分别说明。

Phase 9 记录节选

项目均衡检索 · 对照实验

目标项目出现增加,有效证据未变;不复用 Phase 8 的人工语义评分。

阅读记录节选
## Phase 9 protocol — project-balanced comparison retrieval

### Hypothesis

For cross-project questions, retrieving a larger global semantic candidate pool and selecting the final context across project groups will improve relevant Project Evidence Coverage without changing Standard Search or weakening insufficient-evidence refusal.

### Controlled change

| Variable | Standard baseline | Compare iteration |
|---|---|---|
| Query embedding | Voyage, `input_type: "query"` | Same |
| Initial retrieval | Global Top-5 | Global Top-20 candidates |
| Final context | 5 chunks | 5 chunks, deterministic project-balanced selection |
| Claude model/prompt | Existing grounded-generation path | Same |
| Citation/retry behavior | Existing path | Same |
| Stored document embeddings | Existing vectors | Same — no re-embedding |

The main experimental variable is source selection. Compare mode does not add metadata filtering, reranking, hybrid search, a similarity threshold, or intent classification.

### Coverage metrics

- **Project Presence Coverage**: target projects represented by at least one selected chunk / target projects expected. This is mechanical and can be calculated from `project_id`.
- **Project Evidence Coverage**: target projects with at least one selected chunk that actually contains evidence needed for the comparison / target projects expected. Presence alone does not count. In this privacy-preserving run, expected evidence was checked in memory with deterministic patterns; it is a narrower proxy than a claim-by-claim human review.

### Test suite

- Standard regression: K1, K4, K6, K7; verify Hit@3 remains 4/4.
- Comparison A/B: C1–C4 plus new C5; run each in Standard and Compare modes.
- Unknown-answer regression: U1 and U4; run each in Standard and Compare modes.
- For comparison answers, record evidence coverage, answer relevance, groundedness, semantic citation accuracy, no-answer behavior, end-to-end latency, and actual Voyage/Claude call count.

Planned normal-path usage: 18 Voyage query-embedding calls and 14 Claude generation calls. Automatic bounded retries are counted if they occur. No document embedding is regenerated.

### Results (run 2026-09-02)

The hypothesis was **not supported strongly enough to promote project balancing as the new retrieval default**. Compare improved mechanical target-project presence in C1 and C2, but did not improve the expected-evidence count for any comparison. The candidate pool also contained four embedded project groups because a historical synthetic smoke-test project remains in the local knowledge base; blindly balancing every `project_id` therefore spent one of five context slots on a non-target project.

| Metric | Standard | Compare | Interpretation |
|---|---:|---:|---|
| Known-answer Hit@3 (K1, K4, K6, K7) | 4/4 | Not run | No regression; Standard is the unchanged Phase 8 path |
| Target Project Presence Coverage, C1–C4 | 9/11 | 11/11 | Mechanical coverage improved, but Compare represented 4 total project groups rather than only the 3 evaluation targets |
| Project Evidence Coverage, C1–C4 | 6/11 | 6/11 | No target-project semantic-evidence improvement from balancing |
| Positive-evidence opportunity check, C1–C4 | 6/9 | 6/9 | C1 has one positive method claim; absence in the other projects cannot be proven from one retrieved chunk each |
| C4 evidence coverage | 1/3; `no_evidence` | 1/3; `no_evidence` | The original area-comparison miss remains; refusal stayed correct |
| Answer Relevance | 3 successful answered comparisons; no independent numerical regrade | 4 successful answered comparisons; no independent numerical regrade | Raw answers were not logged, so Phase 8's manual score is not reused as a Phase 9 score |
| Groundedness / semantic Citation Accuracy proxy | Expected cited evidence present in 3/3 answered comparisons; citation indices valid | Expected cited evidence present in 4/4 answered comparisons; citation indices valid | Useful evidence check, but not a claim-by-claim semantic certification |
| Unknown-answer refusal (U1, U4) | 2/2 | 2/2 | No-answer behaviour did not regress |
| Successful comparison-answer latency | median 6.9s, range 6.5–8.3s (n=4) | median 6.7s, range 4.7–9.3s (n=5) | No meaningful latency penalty was observed at this tiny sample size |
| Formal experiment API calls | 11 Voyage, 8 Claude attempts | 7 Voyage, 7 Claude calls | Total 18 Voyage and 15 Claude; Standard C5 used both D18 attempts and still failed |

C5 was the new “cover every project” question. Compare returned an answer with 3/3 target-project presence but only 2/3 expected positive-evidence coverage. Standard failed structured generation twice within the initial request; one separate targeted rerun also failed after both D18 attempts. These failures are preserved as observed rather than replaced with the successful Compare result. They show that bounded retry mitigates a failure category but cannot guarantee successful structured output.

All successful answered responses had valid citation indices. A privacy-preserving deterministic check also confirmed that cited sources contained the expected positive-evidence patterns counted above. This is **not equivalent to Phase 8's manual, claim-by-claim semantic review**: raw answers and excerpts were intentionally not logged or printed during this Codex-run evaluation. Therefore Phase 9 does not invent new numerical claims for full Answer Relevance, Groundedness, or semantic Citation Accuracy. The evidence available supports “expected cited evidence present and citation indices valid,” not “every generated claim independently certified.”

### Product decision from the experiment

- Keep Standard Search as the default and do not change its retrieval logic.
- Keep Compare Projects explicitly labelled **Experimental** so the A/B behaviour remains inspectable, but do not claim that it fixes C4 or improves semantic evidence coverage.
- Do not add an LLM intent classifier, hybrid search, reranker, metadata filtering, or new embedding model in Phase 9.
- A future target-project picker or metadata-scoped retrieval could prevent non-target groups from consuming context, but that is a new product decision rather than a hidden extension of this experiment.

### Actual provider usage note

The valid formal batch used 18 Voyage query embeddings and 15 Claude generation attempts. Targeted verification added 5 Voyage calls and 2 Claude attempts. Three additional Voyage requests succeeded during rate-limit/setup diagnostics; requests rejected with Voyage 429 and the sandboxed production-server failures did not produce usable results. Total successful/provider-accepted activity for this work session was therefore 26 Voyage calls and 17 Claude attempts. No document embedding was regenerated.
下载记录 · Markdown ↓

归档日期:2026.09.07。下载文件附来源位置与 SHA-256,便于比对;哈希说明文件对应关系,不构成独立验证。