TAB-DR-003 · Project Diagnosis · 2026-09-07

TAB 2단계 전략 — 직원 채택과 실용화

1단계(시스템 구축)는 성공했지만 직원 채택은 8월 19일 이후 멈췄다. 질문 30건을 전수 조사해 원인을 규명하고, 최우선 병목이던 봇의 지식 접근 제약을 같은 날 제거했다. 진단과 조치를 함께 담은 첫 처방형 리포트.

종합 스코어카드 Scorecard

6개 평가 축의 가중 평균. 괄호는 DR-002(87점, 3일 전) 대비 변화.

지식 자산의 깊이·품질 자산 규모는 3일간 변동 없음 — 다만 봇이 실제로 접근하는 범위는 4.1배로 확대 (→)
92 강점
데이터 정확성 & 거버넌스 답변보기 31개 문서 카테고리·번역 결손 0건 정비 (→)
92 강점
라이브 시스템 & 접근성 search_wiki(55→225개) + 사용현황·온보딩·주간동향 3개 섹션 신설 (↑5)
95 강점
시스템 안정성 & 운영 변동 없음 — 랩탑 단일의존 구조 그대로 (→)
82 양호
실행 연결성 & 직원 채택 DR-002의 80점은 과대평가(9월 0건·신규 0·실패율 60%). 측정·온보딩 인프라 확보로 일부 회복 (↓7)
73 보완 필요
보안 & 접근제어 Power BI 공개 링크 제거로 3회차 이월 갭 해소. RBAC·staging 노출은 미해결 (↑8)
84 양호
채택 축을 내린 이유 — 정직한 정정. 종합은 87로 같지만 채택 축을 80→73으로 내렸다. 시스템이 나빠져서가 아니라, DR-002가 채택 축에 준 80점의 근거가 부족했음이 전수 조사로 드러났기 때문이다. DR-002는 4명의 실사용 사례를 근거로 삼았지만, 30건 전수 분석 결과 9월 질문 0건 · 8/19 이후 신규 유입 사실상 0 · 답변 실패율 60%가 확인됐다. 사례는 있었으나 지속·확산은 없었다. 점수를 지키는 것보다 정정하는 것이 이 시리즈의 가치를 지킨다.

무엇을 발견했나 Findings

질문 로그 30건, 컨텍스트 번들 코드, 부서 폴더를 전수 조사한 결과.

📉 채택 정지

  • 8/19 정점 하루 6명·7건
  • 9월 1~7일 질문 0건
  • 8/19 이후 신규 유입 사실상 0
  • 첫 사용자 Han·jin은 이후 기록 없음
  • "쓰다가 줄었다"가 아니라 "한 번 써보고 안 돌아왔다"

❌ 실패율 60%

  • 30건 중 18건이 "근거 없음" 포함
  • 그런데 답은 볼트 안에 있었다
  • 봇 스스로: "존재한다는 언급은 있으나 실제 내용은 컨텍스트에 없다"
  • 지식의 부재가 아니라 접근의 부재

🔒 55개 파일의 벽

  • 봇의 뇌 = 55개 파일 583KB
  • Raw Sources 168개(1.4MB) 전부 비가시
  • 제품 마스터스펙·가격표·서비스 매뉴얼 58건·조직도 모두 해당
  • 매 질문 ≈239K 토큰 전량 주입 → 더 늘릴 수 없는 구조

트랙 A 완료 — 병목 제거 Track A Shipped

진단 당일 구현·검증·프로덕션 배포까지 완료했다.

A1 검색
봇이 접근하는 문서를 55개 → 225개로 넓혔다 (4.1배).볼트 전체를 KV에 색인하고 search_wiki·read_wiki_doc 도구로 필요할 때 검색·열람한다. 전량 주입을 키우지 않은 이유는 583KB가 이미 매 질문 ≈239K 토큰이라 1.4MB를 더하면 컨텍스트를 초과하고 관련 문서 2~3개가 200개 사이에 묻히기 때문. 변경분만 업로드하므로 매시간 예약 실행에 부담이 없다.
A2 검증
회귀 테스트 18/18 통과 — "느낌"이 아니라 숫자로 관리한다.실제 실패 질문 16건 + "볼트에 없어야 정상"인 지식공백 2건을 케이스화. 후자가 초기 스코어링에서 "정책" 한 단어로 8건을 반환하던 문제를 잡아냈다. Han이 근거 없음을 받았던 질문이 이제 제품 마스터스펙을 1순위로 찾아낸다.
A3 정직성
"근거 없음"이 막다른 길에서 수집 요청으로 바뀐다.근거 없음 선언 전 볼트 검색 최소 1회 필수. 그래도 없으면 무엇을 찾아봤는지·어떤 자료가 필요한지·지금 대신 확인할 경로를 답하고, 필요 자료가 관리자에게 자동 메일로 전달된다. 실패가 볼트를 키우는 입력이 된다.
A4 측정
답변마다 👍/👎 — 품질을 주 단위로 볼 수 있게 됐다.👎는 한 줄 사유를 받아 답변 레코드에 기록된다. 지금까지 답변 품질은 손으로 문서를 읽어야만 알 수 있었다.
부수
답변보기 정비 — 31개 문서 결손 0건."시스템·IT" 주제 신설, 카테고리 미지정 21건 전부 분류해 "기타" 표시 0건. 질문 미리보기 번역 19건·본문 번역 2건 신규 작성으로 카테고리·제목·질문·본문 번역 결손을 모두 없앴다.
유보
봇의 종단 응답 품질은 아직 미검증이다.staging에 API 키 시크릿이 없고 프로덕션은 Access 뒤에 있어 자동 검증이 막힌다. 검증 범위는 검색·문서조회 경로까지(실제 KV 인덱스 기준 6/6, 로컬 18/18). 축 3의 +3은 "접근 가능 지식이 4.1배로 늘고 그 경로가 실측 검증됐다"는 근거이지 "답변이 좋아졌다"는 검증이 아니다.

트랙 B 완료 + Power BI 제거 + 주간 동향 Same Day, Later

트랙 A 배포 후 같은 날 오후, 트랙 B 전체와 3회차 이월 보안 갭, 트랙 C의 첫 항목까지 처리했다.

B1 식별
이름 13가지 표기 → 8명. 측정이 가능해졌다.정본 명부(people.json)를 두고 입력 경로 자체를 고쳤다 — 로그인 이메일로 본인을 자동 선택하게 해, 이름을 타이핑하던 구조가 만들던 분산을 없앴다. 과거 기록은 렌더 시점에만 통합하고 볼트 원본은 수정하지 않았다.
B2 측정
이 리포트가 손으로 센 숫자를 이제 매 빌드마다 자동 집계한다./usage — 누적 질문·고유 사용자·실패율·월별 추이·사용자별·주제별·답하지 못한 질문 목록. 월별 차트는 빈 달도 표시한다 — 9월 공백이 바로 이 대시보드가 드러내야 할 신호이기 때문이다.
B3 온보딩
DR-001부터 3회차 미착수였던 온보딩 페이지를 만들었다./start — 영업·서비스/부품·물류/재고·경영 4개 탭에 실제 질문 예시 26개. 트랙 A로 새로 답할 수 있게 된 8개는 NEW 배지로 구분. 지금까지 모든 사용자가 /ask를 스스로 발견했지 안내를 받은 게 아니었다.
보안
3회차 이월된 최대 보안 갭이 닫혔다 — 전환이 아니라 제거로.완전공개 "웹에 게시" 링크를 URL 비공개에만 의존해 노출하던 Power BI 페이지를 사이트에서 완전히 제거했다(TAB·K-Master 대시보드가 같은 보고를 대체). 더 이상 지킬 필요가 없는 자산을 지키려 애쓰는 대신 없앤 쪽이 옳았다.
트랙 C
주간 동향 보고서가 매주 월요일 자동 발행된다./digest — 실적은 KPI 규칙을 코드에 고정한 스크립트가 집계하고(에이전트가 재계산하지 않는다), 마켓스캔은 7일 창으로 매주 새로 돌린다. 에이전트의 주간 작업은 마크다운 1개 쓰기뿐이고 렌더·배포는 자동이다. 1호 발행 완료.
유보
UI 버그 3건이 프로덕션에 나갔다가 당일 수정됐다.셋 다 새 페이지를 만들며 기존 공유 자산의 계약을 확인하지 않은 것이 원인이다(공유 CSS 이중 포장, flex 익명 항목, 전역 클래스 충돌). 세 번 모두 배포 전 자동 검사는 통과했다 — 값이 HTML에 있다는 것과 화면에 그려진다는 것은 다른 문제이며, UI 변경은 사람 확인이 필요하다는 것이 이번 회차의 실전 교훈이다.

남은 리스크 Remaining Risks

측정 불가
사용자 이름 표기가 갈려 사용량을 셀 수 없다."Nathan 매니저/매니져", "Siwoo Lee/siwoo/Siwoo", "KEVIN/Kevin 사장님" — 트랙 A의 효과를 측정하려면 로그인 이메일 기준 표준화가 먼저다.
내부 지식
13개 부서 폴더 중 7개가 여전히 비어 있다.재무·운영물류·인사·고객CRM·프로세스SOP·전략기획·컴플라이언스가 문서 0건. 직원 질문은 정확히 이 빈 곳을 향했다 — 검색을 열어도 없는 자료는 찾을 수 없다.
보안 이월
Power BI 공개 링크가 3회차 연속 미착수.DR-001의 즉시 항목이 세 번째 리포트까지 그대로다. RBAC 미구축, staging 도메인 미인증 노출도 여전히 미해결.
산출물 공백
직원이 가장 원한 "산출물 생성"은 아직 기능이 아니다.Sushi Hub 제안서는 1회성 수작업이었다. 보고서·PPT 생성기가 트랙 C의 핵심.
단일 의존
bus factor 1과 랩탑 SPOF가 3회차 연속 동일.런북은 있지만 다른 사람이 실제로 밟아본 적이 없다.

앞으로 해야 할 일 Roadmap

트랙 A는 완료됐다. 다음은 측정 가능하게 만들고(B), 산출물을 만들고(C), 빈 지식을 채우는(D) 순서다.

즉시 0–1개월 · 트랙 B

  1. 사용자 식별 표준화 — 로그인 이메일을 정본 키로. 트랙 A의 효과를 측정하려면 이것이 선행되어야 한다.
  2. 사용량 대시보드 — 질문 수·고유 사용자·부서별 분포·실패율·👍/👎 비율.
  3. 부서별 온보딩 카드 + 팀 시연 (DR-001부터 이월) — 영업/서비스/물류/경영 각각 "이건 물어봐도 된다" 예시 10개.
  4. Power BI Embedded 전환 착수 (3회차 연속 이월) · staging 도메인 Access 편입 (설정 5분).

단기 1–3개월 · 트랙 C

  1. 보고서·발표자료 생성기 — Sushi Hub 패턴의 제품화. 웹 슬라이드(3개국어) + PPTX 출력.
  2. 데이터 셀프서비스 — 자연어 → 표 → CSV/Excel 다운로드.
  3. 딜러 360 뷰 — 매출추이·주문이력·리베이트 잔여·서비스 이력·재고·경쟁사 취급을 한 화면에.
  4. 주간 브리핑 메일(Pull→Push) · 랩탑 장애 자동 알림 · RBAC 설계 착수.

중기 3–6개월 · 트랙 D

  1. 내부 지식 채우기 — 직원 질문이 향한 순서로 09_프로세스_SOP → 08_고객_CRM → 06_운영_물류 → 제품 심화.
  2. 운영 2인 이상 체제 — bus factor 1 해소.
  3. 랩탑 → 상시 서버 완전 이전 — DR-002의 두 차례 장애가 긴급도를 실증했다.
  4. 전략 플레이북 수치 1차 출처 검증 (DR-001부터 이월).

직원 실전 팁 Field Tips

트랙 A로 물어볼 수 있는 범위가 크게 넓어졌다. 예전에 "근거 없음"을 받았던 주제도 다시 시도해볼 가치가 있다.

★ 제품 스펙·단종 모델을 물어봐도 된다

제품 마스터스펙 원본을 이제 봇이 직접 읽는다. 단종 모델, 부품 호환, 냉매 종류까지 확인 가능.

“KUR18-3-N 스펙과 냉매 알려줘”

★ 서비스·부품 매뉴얼도 검색된다

TA Service 매뉴얼 58건이 검색 대상에 들어왔다. 다만 “부품”/“서비스”/“AS”를 명시해야 그 데이터를 본다.

“언더카운터 서비스 매뉴얼에서 컴프레서 교체 절차”

★ 가격표·카탈로그·조직도·가이드북

사내 원본 자료가 검색 범위에 들어왔다. 예전에 “근거 없음”이던 질문을 다시 해보라.

“내부 가이드북의 리턴 절차 알려줘”

★ 답변이 아쉬우면 👎를 눌러달라

한 줄 사유만 남기면 된다. 이것이 다음 개선의 우선순위가 된다.

답변 하단 → 👎 → 한 줄

★ “근거 없음”을 받아도 닫지 말 것

이제 봇이 어떤 자료가 있으면 답할 수 있는지 함께 알려주고, 그 요청이 관리자에게 자동 전달된다.

답변에 적힌 “필요한 자료”를 확인

매출·실적은 시트 뒤지지 말고 물어보라

월별 매출, 딜러 실적, 전년 대비는 /ask가 시트를 직접 조회해 답한다.

“2026년 8월 매출 얼마야?”

K-Master는 이름을 명시하라

명시하지 않으면 기본은 Turbo Air 본브랜드다. 두 브랜드 데이터는 절대 합산되지 않는다.

“K-Master 이번달 매출은?”

중요한 결정엔 한 번 더 확인

/ask는 근거가 약하면 스스로 “불확실”이라 말한다. 큰 숫자는 원본으로 교차확인.

“이 숫자 근거 문서 알려줘”
총평. 이번 회차의 핵심은 점수가 아니라 진단의 정확도와 즉시 실행이다. 같은 날 트랙 A에 이어 트랙 B 전체(사용자 식별·사용현황·온보딩)와 3회차 이월 보안 갭(Power BI 공개 링크) 제거, 주간 동향 자동 발행까지 마쳤다. "직원이 안 쓴다"는 막연한 불안을 30건 전수 조사로 측정 가능한 사실(9월 0건 · 실패율 60% · 55개 파일 병목)로 바꿨고, 최우선 병목을 같은 날 제거해 봇의 지식 접근을 4.1배로 넓혔다. 그러나 기술적 병목이 사라졌다고 채택이 돌아오지는 않는다. 다음 관문은 코드가 아니라 측정(트랙 B)과 산출물(트랙 C)이며, 무엇보다 먼저 사용자 식별을 통일해 개선 효과를 실제로 셀 수 있어야 한다.
Turbo Air Brain (TAB) 프로젝트 · 내부 진단 리포트TAB-DR-003 · 2026-09-07 · 87/100 · A-

TAB-DR-003 · 项目诊断 · 2026-09-07

TAB 第二阶段战略 —— 员工采用与实用化

第一阶段(系统建设)取得成功,但员工采用自 8 月 19 日之后停滞。通过对 30 个提问的全量调查查明原因,并在同一天清除了最优先的瓶颈——机器人的知识访问限制。这是首份将诊断与处置合并呈现的报告。

综合评分卡 Scorecard

6 个评估维度的加权平均。括号为较 DR-002(87 分,3 天前)的变化。

知识资产的深度·质量 3 天内资产规模无变化 —— 但机器人实际可访问的范围扩大至 4.1 倍 (→)
92 优势
数据准确性 & 治理 「查看答案」31 份文档的分类·翻译缺失已全部补齐为 0 件 (→)
92 优势
实时系统 & 可访问性 search_wiki(55→225 份)+ 新增使用情况·入门·每周动态三个板块 (↑5)
95 优势
系统稳定性 & 运营 无变化 —— 笔记本单点依赖结构依旧 (→)
82 良好
执行联结性 & 员工采用 DR-002 的 80 分属高估(9 月 0 件·新用户 0·失败率 60%)。因测量与入门基础设施到位而部分回升 (↓7)
73 需要补强
安全 & 访问控制 移除 Power BI 公开链接,化解连续三期遗留的缺口。RBAC 与 staging 暴露仍未解决 (↑8)
84 良好
下调采用维度的原因 —— 诚实的修正。 综合仍为 87,但采用维度从 80 下调至 73。并非系统变差,而是全量调查揭示出 DR-002 给采用维度的 80 分依据不足。DR-002 以 4 位用户的实际使用案例为依据,但 30 件全量分析显示:9 月提问 0 件、8 月 19 日后新用户实际为 0、回答失败率 60%。案例存在,但没有持续与扩散。相比维持分数,修正分数才能守住本系列的价值。

发现了什么 Findings

对 30 件提问日志、上下文捆绑代码、部门文件夹进行全量调查的结果。

📉 采用停滞

  • 8/19 峰值 单日 6 人·7 件
  • 9 月 1~7 日提问 0 件
  • 8/19 之后新用户实际为 0
  • 首批用户 Han·jin 此后无记录
  • 不是「用着用着变少」,而是「试用一次后再没回来」

❌ 失败率 60%

  • 30 件中 18 件含「无依据」
  • 然而答案本就在保管库里
  • 机器人自述:「有提到其存在,但实际内容不在上下文中」
  • 不是知识缺失,而是访问缺失

🔒 55 个文件的墙

  • 机器人的大脑 = 55 个文件 583KB
  • Raw Sources 168 份(1.4MB)全部不可见
  • 产品主规格·价格表·58 份服务手册·组织图均在其中
  • 每次提问全量注入约 239K tokens → 无法再扩展的结构

轨道 A 完成 —— 清除瓶颈 Track A Shipped

在诊断当天完成了实现、验证与生产部署。

A1 检索
机器人可访问的文档从 55 份扩大到 225 份(4.1 倍)。将保管库全量索引至 KV,通过 search_wiki·read_wiki_doc 工具按需检索与阅读。未选择扩大全量注入,是因为 583KB 已相当于每次提问约 239K tokens,再加 1.4MB 会超出上下文窗口,且 2~3 份相关文档会淹没在 200 份之中。仅上传变更部分,因此对每小时的预约执行没有负担。
A2 验证
回归测试 18/18 通过 —— 用数字而非「感觉」来管理。将 16 个真实失败提问 + 2 个「保管库中本就不该有」的知识空白做成用例。后者揪出了初期评分中仅凭「政策」一词就返回 8 件的问题。Han 曾收到「无依据」的提问,如今能将产品主规格排在第 1 位。
A3 诚实性
「无依据」从死胡同变为收集请求。宣告无依据前必须至少检索保管库一次。若仍无结果,则说明查了什么、需要哪些资料、当下可通过什么途径确认,并将所需资料自动邮件发送给管理员。失败因而成为壮大保管库的输入。
A4 测量
每条回答附 👍/👎 —— 质量可以按周查看了。👎 会附带一行理由记录到回答档案中。此前回答质量只能靠人工逐份阅读才能得知。
附带
「查看答案」整备 —— 31 份文档缺失 0 件。新增「系统·IT」主题,为 21 份未分类文档全部归类,「其他」显示降为 0 件。新增 19 条提问预览翻译与 2 份正文翻译,分类·标题·提问·正文翻译的缺失全部清零。
保留
机器人的端到端回答质量尚未验证。staging 没有 API 密钥,生产环境位于 Access 之后,自动验证受阻。验证范围止于检索与文档读取路径(基于真实 KV 索引 6/6,本地 18/18)。维度 3 的 +3 依据是「可访问知识扩大 4.1 倍且该路径已实测验证」,而非「回答质量已改善」的验证。

轨道 B 完成 + 移除 Power BI + 每周动态 Same Day, Later

轨道 A 部署后的同一天下午,完成了轨道 B 全部、连续三期遗留的安全缺口,以及轨道 C 的第一个项目。

B1 标识
13 种姓名写法 → 8 人。测量成为可能。建立正式名册(people.json),并修正了输入路径本身 —— 通过登录邮箱自动选中本人,消除了手动输入造成写法分散的结构。历史记录仅在渲染时统一,保管库原件不作修改。
B2 测量
本报告靠人工统计的数字,现在每次构建自动汇总。/usage —— 累计提问·独立用户·失败率·每月趋势·按用户·按主题·未能回答的问题清单。月度图表连空白月份也会显示 —— 9 月的空白正是这个仪表盘应当揭示的信号。
B3 入门
自 DR-001 起连续三期未启动的入门页面已建成。/start —— 销售·服务/配件·物流/库存·经营四个标签页,共 26 个真实提问示例,其中因轨道 A 而新可回答的 8 个标注 NEW。此前所有用户都是自行发现 /ask,而非经过引导。
安全
连续三期遗留的最大安全缺口已关闭 —— 不是迁移,而是移除。仅靠链接不公开来保护的完全公开 Power BI 页面已从站点彻底移除(TAB·K-Master 仪表盘可替代同样的报表)。与其费力守护一个不再需要守护的资产,不如直接去掉。
轨道 C
每周动态报告将于每周一自动发布。/digest —— 业绩由把 KPI 规则固化在代码中的脚本汇总(智能体不重新计算),市场扫描每周以 7 天窗口重新执行。智能体每周的工作只是写一个 markdown 文件,渲染与部署均自动完成。第 1 期已发布。
保留
3 个 UI 缺陷曾进入生产环境,均于当天修复。三者根源相同:新建页面时未确认既有共享资产的契约(共享 CSS 被二次包裹、flex 匿名项、全局类名冲突)。三次的部署前自动检查都通过了 —— 值存在于 HTML 与正确渲染到屏幕是两回事,UI 变更需要人工确认,这是本期的实战教训。

遗留风险 Remaining Risks

无法测量
用户姓名写法不一,导致使用量无法统计。「Nathan 매니저/매니져」「Siwoo Lee/siwoo/Siwoo」「KEVIN/Kevin 사장님」—— 要衡量轨道 A 的成效,必须先以登录邮箱为准统一标识。
内部知识
13 个部门文件夹中仍有 7 个是空的。财务·运营物流·人事·客户 CRM·流程 SOP·战略企划·合规均为 0 份文档。员工提问恰恰指向这些空白 —— 即便打开检索,没有的资料也找不到。
安全遗留
Power BI 公开链接连续三期未启动。DR-001 的「立即」事项到第三份报告依旧原样。RBAC 未建、staging 域名未认证暴露也仍未解决。
成果空白
员工最想要的「生成成果」尚未成为功能。Sushi Hub 提案书是一次性手工作业。报告·PPT 生成器是轨道 C 的核心。
单点依赖
bus factor 1 与笔记本 SPOF 连续三期不变。手册虽在,但从未有他人真正按其操作过。

后续待办事项 Roadmap

轨道 A 已完成。接下来的顺序是使其可测量(B)、产出成果(C)、填补空白知识(D)。

立即 0–1 个月 · 轨道 B

  1. 统一用户标识 —— 以登录邮箱为正式主键。要衡量轨道 A 的成效,这一步必须先行。
  2. 使用量仪表盘 —— 提问数·独立用户·部门分布·失败率·👍/👎 比例。
  3. 分部门入门卡片 + 团队演示 (自 DR-001 遗留) —— 为销售/服务/物流/管理层各准备 10 个「可以这样问」的示例。
  4. 启动 Power BI Embedded 转换 (连续三期遗留) · 将 staging 域名纳入 Access (设置仅需 5 分钟)。

短期 1–3 个月 · 轨道 C

  1. 报告·演示资料生成器 —— 将 Sushi Hub 模式产品化。输出网页幻灯片(三语)+ PPTX。
  2. 数据自助服务 —— 自然语言 → 表格 → CSV/Excel 下载。
  3. 经销商 360 视图 —— 销售趋势·订单历史·返利余额·服务履历·库存·所经营竞品集中于一屏。
  4. 每周简报邮件(Pull→Push) · 笔记本故障自动告警 · 启动 RBAC 设计。

中期 3–6 个月 · 轨道 D

  1. 填补内部知识 —— 按员工提问指向的顺序:09_프로세스_SOP → 08_고객_CRM → 06_운영_물류 → 产品深化。
  2. 双人以上运营体制 —— 消除 bus factor 1。
  3. 笔记本 → 常驻服务器完全迁移 —— DR-002 的两次故障已印证其紧迫性。
  4. 验证战略手册数值的一手来源 (自 DR-001 遗留)。

员工实用技巧 Field Tips

通过轨道 A,可提问的范围大幅拓宽。以前收到「无依据」的主题也值得再试一次。

★ 可以询问产品规格·停产机型

机器人现在能直接阅读产品主规格原件。停产机型、配件兼容、制冷剂类型均可确认。

"告诉我 KUR18-3-N 的规格与制冷剂"

★ 服务·配件手册也可检索

58 份 TA Service 手册已纳入检索范围。但需在提问中写明「配件」「服务」「AS」才会查阅该数据。

"Undercounter 服务手册中的压缩机更换步骤"

★ 价格表·产品目录·组织图·指南手册

公司内部原始资料已进入检索范围。请把以前得到「无依据」的问题再问一次。

"内部指南手册中的退货流程"

★ 回答不满意请点 👎

只需留下一行理由,这将成为下一步改进的优先级依据。

回答下方 → 👎 → 一行说明

★ 收到「无依据」也不要直接关闭

机器人现在会一并告知需要哪些资料才能作答,该请求会自动转达给管理员。

查看回答中写明的「所需资料」

销售·业绩不用翻表格,直接问

月度销售额、经销商业绩、同比变化,/ask 会直接查询表格作答。

"2026 年 8 月销售额是多少?"

K-Master 请写明名称

不写明则默认为 Turbo Air 主品牌。两个品牌的数据绝不合并计算。

"K-Master 本月销售额是多少?"

重要决策请再确认一次

/ask 在依据不足时会主动说明「不确定」。重大数字请与原件交叉核实。

"告诉我这个数字的依据文档"
总评。 本期的核心不是分数,而是诊断的精确度与即时执行。同一天在轨道 A 之后,还完成了轨道 B 全部(用户标识·使用情况·入门引导)、移除连续三期遗留的安全缺口(Power BI 公开链接),以及每周动态的自动发布。「员工不用」这一模糊的不安,通过 30 件全量调查被转化为可测量的事实(9 月 0 件 · 失败率 60% · 55 个文件的瓶颈),并在同一天清除了最优先的瓶颈,将机器人的知识访问扩大至 4.1 倍。然而技术瓶颈消失并不等于采用会自动回来。下一个关卡不在代码,而在测量(轨道 B)与成果(轨道 C);尤其必须先统一用户标识,才能真正数得清改进的成效。
Turbo Air Brain (TAB) 项目 · 内部诊断报告TAB-DR-003 · 2026-09-07 · 87/100 · A-

TAB-DR-003 · Project Diagnosis · 2026-09-07

TAB Phase 2 Strategy — Adoption and Practical Value

Phase one (building the system) succeeded, but staff adoption stopped after 19 August. A full read of all 30 logged questions identified the cause, and the top bottleneck — the bot's restricted access to the vault — was removed the same day. The first report in this series to carry both the diagnosis and the treatment.

Overall Scorecard Scorecard

A weighted average across 6 axes. Parentheses show change vs. DR-002 (87, three days ago).

Depth & Quality of Knowledge Assets No change in the assets themselves over three days — but the share the bot can actually reach grew 4.1x (→)
92 Strength
Data Accuracy & Governance All 31 Answers documents brought to zero gaps in category and translation (→)
92 Strength
Live System & Accessibility search_wiki (55→225 documents) plus three new sections: usage, onboarding and the weekly digest (↑5)
95 Strength
System Stability & Operations Unchanged — the single-laptop dependency remains (→)
82 Good
Execution Connectivity & Staff Adoption DR-002's 80 was an overstatement (zero in September, no new users, 60% failure). Partly recovered as measurement and onboarding landed (↓7)
73 Needs Work
Security & Access Control Removing the public Power BI link closed a gap carried for three reports. RBAC and the staging exposure remain (↑8)
84 Good
Why the adoption axis came down — an honest correction. The overall score held at 87, but adoption was cut from 80 to 73 - not because the system got worse, but because a full audit showed DR-002's 80 on the adoption axis was poorly evidenced. DR-002 relied on anecdotes of four people using it; the full read of 30 questions found zero questions in September, essentially no new users after 19 August, and a 60% answer-failure rate. The anecdotes were real; the continuation was not. Correcting the score protects this series' value more than defending it would.

What We Found Findings

From a full audit of the 30 logged questions, the context-bundle code, and the department folders.

📉 Adoption stopped

  • Peak on 19 Aug: 6 people, 7 questions in one day
  • Zero questions 1-7 September
  • Essentially no new users since 19 Aug
  • First-time users Han and jin never returned
  • Not "usage tapered off" but "tried once, never came back"

❌ 60% failure rate

  • 18 of 30 contained "no evidence"
  • Yet the answers were in the vault
  • The bot itself: "there is a mention that it exists, but its contents are not in this context"
  • Not missing knowledge — missing access

🔒 The 55-file wall

  • The bot's brain = 55 files, 583KB
  • All 168 raw sources (1.4MB) invisible
  • Master specs, price lists, 58 service manuals, org chart — all of it
  • ≈239K tokens injected per question → a structure that could not grow

Track A Shipped — Bottleneck Cleared Track A Shipped

Built, verified and deployed to production on the same day as the diagnosis.

A1 Retrieval
Documents the bot can reach went from 55 to 225 (4.1x).The whole vault is indexed into KV and reached on demand through the search_wiki and read_wiki_doc tools. Growing the always-on bundle was not an option: 583KB already means ≈239K tokens per question, adding 1.4MB would exceed the context window, and the 2-3 relevant documents would be buried among 200. Only changed documents upload, so the hourly scheduled run carries no extra load.
A2 Verification
18/18 regression tests pass — managed by numbers, not impressions.16 real failed questions plus 2 knowledge gaps that should return nothing. The latter caught an early scoring flaw that returned 8 documents off the single common word "policy". The question Han got "no evidence" on now ranks the product master spec first.
A3 Honesty
"No evidence" turns from a dead end into a collection request.The bot must search the vault at least once before declaring no evidence. If nothing turns up it states what it searched, which material would let it answer, and where to check instead — and that request is emailed to the administrator automatically. Failure becomes an input that grows the vault.
A4 Measurement
👍/👎 on every answer — quality is now visible week to week.A 👎 captures a one-line reason onto the answer record. Until now answer quality could only be judged by reading documents by hand.
Alongside
Answers section tidied — zero gaps across 31 documents.Added a "System & IT" topic and categorised all 21 unclassified documents, taking "Other" to zero. Wrote 19 new question-preview translations and 2 full body translations, closing every category, title, question and body translation gap.
Caveat
The bot's end-to-end answer quality is still unverified.Staging has no API key secret and production sits behind Access, which blocks automated checks. Verification covers the search and document-read paths (6/6 against the real KV index, 18/18 locally). The +3 on axis 3 rests on "reachable knowledge grew 4.1x and that path is measured", not on "answers got better".

Track B, Power BI Removal and the Weekly Digest Same Day, Later

The same afternoon Track A shipped, all of Track B, a security gap carried for three reports, and the first Track C item were completed too.

B1 Identity
13 name spellings became 8 people - measurement is now possible.A canonical roster (people.json) plus a fix to the input path itself: the logged-in person is selected automatically from their Access email, removing the typing that caused the divergence. Historical records are unified at render time only - the vault originals are never rewritten.
B2 Measurement
The numbers this report counted by hand are now aggregated on every build./usage - questions, unique users, failure rate, monthly trend, by person, by topic, and a list of unanswered questions. The monthly chart renders empty months too, because a silent September is exactly the signal this dashboard exists to surface.
B3 Onboarding
The onboarding page unstarted since DR-001 now exists./start - four tabs (sales, service and parts, logistics and stock, management) with 26 real example questions; the 8 newly answerable thanks to Track A carry a NEW badge. Until now every user found /ask on their own rather than being shown it.
Security
The biggest carried-over gap is closed - by removal, not migration.The Power BI page, a fully public "publish to web" link protected by nothing but an unlisted URL, was deleted from the site entirely (the TAB and K-Master dashboards cover the same reporting). Better to remove an asset that no longer needs defending than to keep working at defending it.
Track C
A weekly digest now publishes automatically every Monday./digest - performance figures come from a script with the KPI rules fixed in code (the agent never recomputes them), and the market scan re-runs weekly on a 7-day window. The agent's weekly job is writing one markdown file; rendering and deployment are automatic. Issue 1 is published.
Caveat
Three UI bugs reached production and were fixed the same day.All three had the same root: building a new page without checking the contract of an existing shared asset (double-wrapping the shared CSS, an anonymous flex item, a global class collision). All three passed the pre-deploy automated checks - a value being present in the HTML is not the same as it rendering correctly, and UI changes need human eyes. That is this round's practical lesson.

Remaining Risks Remaining Risks

Unmeasurable
Split name spellings make usage impossible to count."Nathan 매니저/매니져", "Siwoo Lee/siwoo/Siwoo", "KEVIN/Kevin 사장님" — measuring Track A's effect requires standardising on the login email first.
Internal Knowledge
Seven of thirteen department folders are still empty.Finance, operations/logistics, HR, customer CRM, process/SOP, strategy and compliance hold zero documents. Staff questions pointed straight at these gaps — opening up search cannot find material that isn't there.
Security Carryover
The public Power BI link is untouched for a third report running.An "immediate" item from DR-001 is unchanged three reports later. RBAC is unbuilt and the staging domain is still exposed without authentication.
Output Gap
The output generation staff wanted most is still not a feature.The Sushi Hub proposal was one-off manual work. A report and deck generator is the core of Track C.
Single Dependency
Bus factor 1 and the laptop SPOF are unchanged for a third report.The runbook exists, but nobody else has ever walked through it.

What's Next Roadmap

Track A is done. The order from here is make it measurable (B), produce outputs (C), fill the empty knowledge (D).

Immediate 0–1 month · Track B

  1. Standardise user identity — the login email as the canonical key. Measuring Track A's effect depends on this going first.
  2. Usage dashboard — questions, unique users, distribution by department, failure rate, 👍/👎 ratio.
  3. Department onboarding cards + team demo (carried from DR-001) — 10 real "you can ask this" examples each for sales, service, logistics and management.
  4. Begin the Power BI Embedded migration (carried three reports running) · bring the staging domain under Access (a 5-minute setting).

Short-term 1–3 months · Track C

  1. Report and deck generator — productising the Sushi Hub pattern. Web slides (three languages) plus PPTX.
  2. Data self-service — natural language → table → CSV/Excel download.
  3. Dealer 360 view — revenue trend, order history, rebate balance, service history, stock and competitor brands on one screen.
  4. Weekly briefing email (pull → push) · automatic failure alerts · start RBAC design.

Mid-term 3–6 months · Track D

  1. Fill the internal knowledge base — in the order staff questions pointed: 09_프로세스_SOP → 08_고객_CRM → 06_운영_물류 → deeper product material.
  2. Two or more operators — eliminate bus factor 1.
  3. Full migration from laptop to an always-on server — DR-002's two incidents demonstrated the urgency.
  4. Verify the strategy playbook's figures against primary sources (carried from DR-001).

Field Tips for Staff Field Tips

Track A widened what you can ask considerably. Topics that previously came back "no evidence" are worth trying again.

★ You can ask about specs and discontinued models

The bot now reads the product master spec directly — discontinued models, parts compatibility and refrigerant types included.

"Give me the spec and refrigerant for KUR18-3-N"

★ Service and parts manuals are searchable

58 TA Service manuals are now in scope. Say "parts", "service" or "AS" explicitly so it looks at that data.

"Compressor replacement steps in the undercounter service manual"

★ Price lists, catalogues, org chart, guide book

Internal source material is now in the search scope. Re-ask the questions that used to return "no evidence".

"What's the return process in the internal guide book?"

★ Press 👎 if an answer disappoints

One line of reasoning is enough — it sets the priority for the next round of improvements.

Below the answer → 👎 → one line

★ Don't close the tab on "no evidence"

The bot now tells you which material would let it answer, and that request goes to the administrator automatically.

Check the "material needed" line in the answer

Don't dig through spreadsheets — just ask

Monthly revenue, dealer performance and year-on-year change are queried straight from the sheet.

"What was revenue in August 2026?"

Name K-Master explicitly

Without it the default is the core Turbo Air brand. The two brands' data is never combined.

"What was K-Master's revenue this month?"

Double-check for important decisions

/ask says "uncertain" when the evidence is thin. Cross-check big numbers against the source.

"Show me the source document for this number"
Overall assessment. What matters this round is not the score but the precision of the diagnosis and how fast it was acted on. Track A was followed the same day by all of Track B (identity, usage, onboarding), removal of the public Power BI exposure carried for three reports, and an automated weekly digest. A vague worry that "staff aren't using it" became measurable fact through a full audit of 30 questions — zero in September, a 60% failure rate, a 55-file bottleneck — and the top bottleneck was removed the same day, widening the bot's reach 4.1x. But clearing a technical bottleneck does not bring adoption back on its own. The next gate is not code but measurement (Track B) and outputs (Track C) — and above all, unifying user identity so the effect of any improvement can actually be counted.