Agent Skill · product research Agent Skill · 제품 리서치

$persona-product-tester

Synthetic users that actually use your product.

설정된 사용자가 제품을 직접 조작합니다.

Configure a user — their goal, current workflow, expertise, device, time pressure, reading habits, recovery budget, and quitting condition — then let them operate your site, app, or prototype one action at a time. You get a session transcript, a debrief in the user's voice, and a separate analyst report where every finding points back at an observed step.

사용자를 설정하세요. 목표, 현재 업무 방식, 숙련도, 기기, 시간 압박, 읽기 습관, 복구 예산, 포기 조건까지 정하면 그 사용자가 사이트·앱·프로토타입을 한 번에 한 행동씩 실제로 조작합니다. 결과로 세션 기록, 사용자 목소리로 쓴 회고, 그리고 모든 발견이 관찰된 단계를 가리키는 별도의 분석가 보고서를 받습니다.

session.transcript
$persona-product-tester
persona  kr-middle-school-math-teacher
scenario exam-bank--assemble-a-quiz
budget   12 actions / 2 recoveries

step 02 saw "문항 담기" chip -> clicked
        result cart badge unchanged
step 05 scanned toolbar for undo -> none visible
step 07 retried from list view -> same result
        recovery budget spent

exit     GAVE_UP  (7 / 12 actions)
report   product-feedback/2026-08-05/...

findings
  OBSERVED          add-to-cart gives no feedback
  PERSONA-INFERRED  skims toolbars, missed menu
  ASSUMPTION        account already had a class
  UNTESTED          export and print paths
$persona-product-tester
페르소나 kr-middle-school-math-teacher
시나리오 exam-bank--assemble-a-quiz
예산     행동 12회 / 복구 2회

단계 02 "문항 담기" 칩이 보임 -> 클릭
        결과 담긴 개수 표시 변화 없음
단계 05 되돌리기를 찾아 툴바를 훑음 -> 안 보임
단계 07 목록 화면에서 재시도 -> 동일
        복구 예산 소진

종료     GAVE_UP  (행동 7 / 12)
보고서   product-feedback/2026-08-05/...

발견
  OBSERVED          담기 동작에 피드백이 없음
  PERSONA-INFERRED  툴바를 훑는 습관 탓에 메뉴를 놓침
  ASSUMPTION        계정에 반이 이미 있다고 가정
  UNTESTED          내보내기·인쇄 경로
Why it is different 무엇이 다른가

Not roleplay. A fixed procedure.

역할극이 아니라 고정된 절차입니다.

Do not imitate how the configured person talks. Let their goal, current method, knowledge, environment, and patience limit change how they use the product. The aim is not the same conclusion — it is the same inspection procedure.

설정된 사람의 말투만 흉내 내지 않습니다. 그 사람의 목표·현재 방식·지식·환경·인내 한계가 제품 사용 행동을 바꾸게 합니다. 목표는 같은 결론이 아니라 같은 검사 절차입니다.

01Demographics are not behavior인구통계는 행동이 아니다

Being in your sixties is not a reason to assign low digital literacy. Domain expertise and digital literacy are tracked as separate fields, and any behavior built from demographics alone is removed from the contract.

60대라는 이유로 디지털 숙련도를 낮게 잡지 않습니다. 업무 숙련도와 디지털 숙련도를 별도 필드로 관리하고, 인구통계만으로 만든 행동은 계약에서 제거합니다.

02Real operation comes first실제 사용이 먼저다

Rather than reading code or requirements and guessing at UX, the screen is operated directly wherever possible. If only a semantic tree was available with no visual rendering, the run does not claim to have verified visual usability.

코드나 요구사항을 읽고 UX를 추측하는 대신 가능한 경우 화면을 직접 조작합니다. visual 화면 없이 semantic tree만 사용했다면 시각적 사용성을 검증했다고 말하지 않습니다.

03Giving up is a valid outcome포기가 정상 결과다

The agent does not force a user who has spent their action and recovery budget into a success. A run that ends in GAVE_UP is a result, not a failed test.

행동·복구 예산을 다 쓴 사용자를 에이전트가 억지로 성공시키지 않습니다. GAVE_UP으로 끝난 실행은 실패한 테스트가 아니라 하나의 결과입니다.

04The user and the analyst are separate사용자와 분석가를 분리한다

In user mode, no diagnosing causes or designing fixes like a UX expert. The debrief speaks only from experience; the analyst report is written afterwards, as a separate document.

사용자 모드에서는 UX 전문가처럼 원인을 진단하거나 해법을 설계하지 않습니다. 회고는 경험만 말하고, 분석가 보고서는 그 뒤에 별도 문서로 작성합니다.

The contract 실행 계약

Five execution criteria

다섯 개의 실행 기준

These hold for every mode and every run. They are what keep a persona session from drifting into a plausible-sounding story.

모든 모드와 모든 실행에 적용됩니다. 페르소나 세션이 그럴듯한 이야기로 흘러가지 않게 막는 장치입니다.

Grounded근거화

Significant behavior traces to a persona field or supplied research. Age, region, and gender alone never imply skill level or preference.

중요한 행동은 페르소나 필드나 제공된 조사 자료에 연결합니다. 연령·지역·성별만으로 숙련도나 선호를 추정하지 않습니다.

Visible가시성

The next action comes from user-visible cues only. Hidden code, requirements, the DOM, and network logs are never used as an answer key.

사용자에게 보이는 단서만으로 다음 행동을 고릅니다. 숨은 코드, 요구사항, DOM, 네트워크 로그를 정답지처럼 사용하지 않습니다.

Finite유한성

Time, action count, retries, memory, and help-seeking are all budgeted. On reaching the contracted stop condition, the persona may quit like a real user.

시간, 행동 수, 재시도, 기억과 도움 요청을 무제한으로 쓰지 않습니다. 계약된 중단 조건에 도달하면 실제 사용자처럼 포기할 수 있습니다.

Isolated격리

Multiple personas run independently, unaware of each other's findings. A retest does not use prior results as an exploration hint.

여러 페르소나는 서로의 발견을 모른 채 독립적으로 실행합니다. 재테스트도 이전 결과를 탐색 힌트로 사용하지 않습니다.

Flagged as synthetic합성표시

Results are a synthetic test signal, not user evidence. They never claim proven demand, purchase intent, real emotion, or statistical frequency.

결과는 사용자 증거가 아니라 합성 테스트 신호입니다. 실제 수요·구매 의사·감정·통계적 빈도를 입증했다고 표현하지 않습니다.

Modes 모드

Five modes, inferred from the request

요청에서 추론하는 다섯 모드

You do not have to name a mode. If a blocking value is missing — the product target, the task, or the auth state — it is asked for once, as a single batch. Everything else falls back to a neutral default recorded as an ASSUMPTION.

모드를 직접 지정하지 않아도 됩니다. 실행을 막는 값인 제품 대상, 과업, 인증 상태가 없으면 한 번에 묶어 요청합니다. 나머지 누락값은 중립 기본값을 쓰고 ASSUMPTION으로 기록합니다.

Mode모드 What it does하는 일
configure Builds a reusable persona YAML plus a readable review card.재사용 가능한 페르소나 YAML과 검토용 페르소나 카드를 만듭니다.
test One persona actually performs one task. This is the default.한 페르소나가 한 과업을 실제로 수행합니다. 기본 모드입니다.
panel Two or more personas run as isolated sessions, then commonalities and differences are merged.2개 이상의 페르소나를 독립 세션으로 실행한 뒤 공통점과 차이를 합칩니다.
retest Compares before and after a fix, holding persona, scenario, and environment constant.같은 페르소나·시나리오·환경으로 수정 전후를 비교합니다.
artifact-review Reviews screens, copy, or a static prototype when real operation is impossible. The limitation is marked throughout the report.실제 조작이 불가능할 때 화면·카피·정적 프로토타입만 검토합니다. 보고서 전 구간에 이 제한을 표시합니다.
Run loop 실행 절차

Seven steps, each with a completion condition

완료 조건이 붙은 일곱 단계

01

Resolve

Read the persona and scenario files, record the source of any real research, and check the required fields from the persona contract. Write the task as "what am I trying to finish", not "evaluate this feature".

페르소나와 시나리오 파일을 읽고 실제 조사 자료의 출처를 기록한 뒤, 페르소나 계약의 필수 필드를 검사합니다. 과업은 "기능을 평가하라"가 아니라 "무엇을 끝내려는가"로 씁니다.

Done when persona requirements, success condition, stop conditions, access method, and safety scope are all stated. 완료 조건: 페르소나 필수값, 과업 성공 조건, 중단 조건, 접근 방식과 안전 범위가 모두 명시되어 있다.
02

Freeze

Pin the behavior contract one line at a time: goal, current method and switching cost, prior knowledge and gaps, reading and navigation style, max actions and recoveries, risk and privacy thresholds, device and time pressure. Never adjust the contract mid-run to fit results.

행동 계약을 한 줄씩 고정합니다. 목표, 현재 방식과 전환 비용, 사전 지식과 모르는 것, 읽기·탐색 성향, 최대 행동 수와 복구 횟수, 위험·개인정보 기준, 기기와 시간 압박. 실행 중 결과에 맞춰 계약을 바꾸지 않습니다.

Done when every significant behavior rule has a source, and assumptions are separated from unknowns. 완료 조건: 모든 중요한 행동 규칙에 출처가 있고, 가정과 미지수가 분리되어 있다.
03

Prepare

Use test accounts and synthetic data. Irreversible actions — purchases, payments, outbound messages, deletions, public posts, production data changes — are never taken without explicit permission. Capture the first screen as evidence.

테스트 계정과 합성 데이터를 사용합니다. 구매·결제·외부 메시지 발송·삭제·공개 게시·운영 데이터 변경 같은 비가역 행동은 명시적 허용 없이 실행하지 않습니다. 첫 화면과 환경을 증거로 남깁니다.

Done when a safe starting point and the permitted action scope are on record. 완료 조건: 안전한 시작점과 허용된 행동 범위가 기록되어 있다.
04

Use

One action at a time. Each action records cue seen → action taken → actual result → short user reaction → state. Skim or read according to the persona's style; do not inspect every label and menu. On confusion, use only the contracted recovery method.

한 번에 한 행동만 합니다. 각 행동마다 보인 단서 → 수행한 행동 → 실제 결과 → 짧은 사용자 반응 → 상태를 기록합니다. 페르소나의 읽기 방식대로 훑거나 읽고, 모든 문구와 메뉴를 완벽히 검사하지 않습니다. 혼란이 생기면 계약된 복구 방식만 사용합니다.

Done when one exit state is reached and every finding links to a real step. 완료 조건: 하나의 종료 상태에 도달했고, 모든 관찰 발견이 실제 단계에 연결되어 있다.
05

Debrief

In the persona's voice, and only this: what they were trying to do, what actually happened, the biggest friction or blocker, and what they would do next. No invented NPS, satisfaction scores, or purchase-intent numbers.

페르소나 목소리로 다음만 말합니다. 하려던 일, 실제로 일어난 일, 가장 큰 불편 또는 막힌 지점, 지금이라면 다음에 할 행동. 임의의 NPS·만족도 점수·구매 의사 점수를 만들지 않습니다.

Done when the debrief rests on experience alone, with no hidden implementation detail or expert judgment. 완료 조건: 회고가 사용 경험에만 근거하고 숨은 구현 정보나 전문가 평가를 포함하지 않는다.
06

Synthesize

Write the analyst report separately, labelling every claim. Severity is rated by task impact, not by taste. Questions about real demand, willingness to pay, genuine emotion, population rates, or legal and accessibility compliance are demoted to hypotheses or handed to real user research.

분석가 보고서를 분리해 작성하고 모든 서술에 라벨을 붙입니다. 심각도는 호불호가 아니라 과업 영향으로 매깁니다. 실제 수요·지불 의사·진짜 감정·모집단 비율·법적·접근성 적합성을 묻는 요청은 가설로 낮추거나 실제 사용자 조사로 넘깁니다.

Done when the report repeatedly marks itself synthetic and every key finding carries evidence, impact, severity, confidence, and the related persona field. 완료 조건: 보고서가 합성 결과임을 반복 표시하고, 모든 핵심 발견에 증거·영향·심각도·신뢰도·관련 페르소나 필드가 있다.
07

Persist

Save to product-feedback/YYYY-MM-DD/<scenario-id>--<persona-id>--<run-id>.md unless the project says otherwise. panel saves both the individual reports and a combined one; retest leads with a table of what was held constant and what changed.

프로젝트에서 다른 위치를 정하지 않았다면 product-feedback/YYYY-MM-DD/<scenario-id>--<persona-id>--<run-id>.md에 저장합니다. panel은 개별 보고서와 종합 보고서를 모두 저장하고, retest는 동일·변경 조건을 먼저 표로 적습니다.

Done when the result carries the persona, scenario, product version, environment, and evidence location needed to reproduce it. 완료 조건: 재현에 필요한 페르소나, 시나리오, 제품 버전, 환경과 증거 위치가 결과에 포함되어 있다.
Exit states & evidence labels 종료 상태와 증거 라벨

Every run ends in exactly one state

모든 실행은 정확히 하나의 상태로 끝납니다

Environment or tooling problems are recorded as a TEST BLOCKER, kept separate from product usability defects — a broken test harness is not a UX finding.

환경 또는 도구 문제는 제품 사용성 결함과 분리해 TEST BLOCKER로 기록합니다. 테스트 도구가 깨진 것은 UX 발견이 아닙니다.

Exit states종료 상태

COMPLETED GAVE_UP BLOCKED SAFETY_STOP

The product is not forced forward when auth or test data is missing, when production could be irreversibly affected, when the product or tooling fails to load, when the persona's budget is spent, or when the success condition is met.

인증이나 테스트 데이터가 없을 때, 운영 환경에 비가역 영향을 줄 위험이 있을 때, 제품이나 테스트 도구가 로드되지 않을 때, 페르소나의 예산이 소진됐을 때, 성공 조건에 도달했을 때는 제품을 억지로 계속 사용하지 않습니다.

Evidence labels증거 라벨

OBSERVED PERSONA-INFERRED ASSUMPTION UNTESTED

OBSERVED — confirmed on a real screen and action. PERSONA-INFERRED — likely because of a configured trait. ASSUMPTION — a neutral placeholder with no evidence, kept so the run could proceed. UNTESTED — never reached or never checked.

OBSERVED — 실제 화면과 행동에서 확인된 사실. PERSONA-INFERRED — 설정된 특성 때문에 가능성이 높다고 해석한 내용. ASSUMPTION — 근거 없이 실행을 위해 둔 중립 가정. UNTESTED — 접근하지 못했거나 확인하지 않은 영역.

Install 설치

Clone it into your skills directory

스킬 디렉터리에 클론하세요

Claude Code (project-local)Claude Code (프로젝트 전용)
mkdir -p .claude/skills
git clone https://github.com/cskwork/persona-product-tester \
  .claude/skills/persona-product-tester
Shared Agent Skills공통 Agent Skills
mkdir -p .agents/skills
git clone https://github.com/cskwork/persona-product-tester \
  .agents/skills/persona-product-tester

Then invoke it

그다음 호출합니다

Claude Code uses /persona-product-tester. The Codex and ChatGPT family use $persona-product-tester, or pick it explicitly in the skill selector. A browser or GUI tool is strongly recommended — without one, the run is limited to artifact-review mode and must state that no real operation happened.

Claude Code는 /persona-product-tester, Codex·ChatGPT 계열은 $persona-product-tester 또는 Skills 선택기에서 명시적으로 선택합니다. 브라우저나 GUI 도구 사용을 강력히 권장합니다. 상호작용 도구가 없으면 artifact-review 모드로 제한하고 실제 조작이 없었음을 명시해야 합니다.

Validate inputs, then run입력을 검증한 뒤 실행
python scripts/validate_inputs.py \
  --persona examples/kr-middle-school-math-teacher.yaml \
  --scenario examples/exam-bank-scenario.yaml

# no PyYAML? prefix with:  uv run --with pyyaml
A real prompt실제 프롬프트
$persona-product-tester
Use personas/my-user.yaml and scenarios/my-task.yaml to actually
operate this local product, and save the result under product-feedback.
$persona-product-tester
personas/my-user.yaml과 scenarios/my-task.yaml로 이 로컬 제품을
실제 조작해 테스트하고 결과를 product-feedback에 저장해.

No configuration yet? Ask for configure mode first and it will build the persona YAML and a review card from assets/persona-template.yaml. 설정이 없다면 configure 모드로 먼저 요청하세요. assets/persona-template.yaml을 바탕으로 페르소나 YAML과 검토용 카드를 만들어 줍니다.

Writing a persona 페르소나 작성

"39-year-old teacher" is not enough

"39세 교사"로는 부족합니다

A useful persona carries information you can act on. When comparing user types, do not just change the age — vary the axes that directly change product use: goal, current method, expertise, permissions, environment.

좋은 페르소나는 행동 가능한 정보를 담습니다. 여러 사용자 유형을 비교할 때는 나이만 바꾸지 말고 목표, 현재 방식, 숙련도, 권한, 환경처럼 제품 사용을 직접 바꾸는 축을 달리하세요.

The job to be done지금 해야 하는 일

The real work they need to finish right now, and what it costs them to fail.

지금 해결해야 하는 실제 일과 실패했을 때의 영향.

Current tooling현재 도구

What they use today and in what order — this is the switching cost your product is up against.

현재 쓰는 도구와 순서. 제품이 상대해야 할 전환 비용입니다.

Two literacies두 종류의 숙련도

Domain expertise and digital literacy, tracked separately. They move independently.

업무 숙련도와 디지털 숙련도를 따로 기록합니다. 둘은 함께 움직이지 않습니다.

Reading style읽기 방식

Do they skim the screen or read it closely? This decides which cues they will ever see.

화면을 훑는지 정독하는지. 어떤 단서를 보게 될지가 여기서 결정됩니다.

Recovery budget복구 예산

How many failures before they seek help or quit, and which actions feel risky enough to avoid.

몇 번 실패하면 도움을 찾거나 포기하는지, 어떤 행동을 위험하다고 느껴 피하는지.

Environment환경

Device, screen size, network, and time pressure — the conditions the session actually runs under.

기기, 화면 크기, 네트워크, 시간 압박. 세션이 실제로 놓이는 조건입니다.

Limits 한계

What this does not do

이 스킬이 하지 않는 일

Do not use it for다음 용도로는 쓰지 마세요

  • Summarizing real customer interviews실제 고객 인터뷰 요약
  • Code review or CSS fixes코드 리뷰나 CSS 수정
  • General design critique일반 디자인 평론
  • Validating market demand or conversion rates on its own시장 수요·구매율 검증만 필요한 경우

This skill is useful for finding early usability defects and for surfacing questions worth taking to real user research. It does not replace real user interviews, usability testing, analytics, or accessibility validation. 이 스킬은 초기 사용성 결함과 실제 사용자 연구 질문을 찾는 데 유용하지만, 실제 사용자 인터뷰·사용성 테스트·분석 데이터·접근성 검증을 대체하지 않습니다.