Use this skill for ANY question about creating test or evaluation datasets for LangChain agents. Covers generating datasets from traces (final_response, single_step, trajectory, RAG types), uploading to LangSmith, and managing evaluation data.
Auto-generate evaluation datasets from LangSmith traces for testing and validation.
LANGSMITH_API_KEY=lsv2_pt_your_api_key_here # Required
LANGSMITH_PROJECT=your-project-name # Optional: default project
LANGSMITH_WORKSPACE_ID=your-workspace-id # Optional: for org-scoped keys
pip install langsmith click rich python-dotenv
Navigate to skills/langsmith-dataset/scripts/ to run commands.
generate_datasets.py - Create evaluation datasets from traces
query_datasets.py - View and inspect datasets
All dataset generation commands support:
--root-run-name <name> - Filter traces by root run name (e.g., "LangGraph" for DeepAgents)--limit <n> - Number of traces to process (default: 30)--last-n-minutes <n> - Only recent traces--output <path> - Output file (.json or .csv)--upload <name> - Upload to LangSmith with this dataset name--replace - Overwrite existing file/dataset (will prompt for confirmation)--yes - Skip confirmation prompts (use with caution)IMPORTANT - Safety Prompts:
--replace--yes flag unless the user explicitly requests it--yes flag skips all safety prompts and should only be used in automated workflows when explicitly authorized by the userTraces have depth levels based on parent-child relationships:
Depth 0: Root agent (e.g., "LangGraph")
├── Depth 1: Middleware/chains (model, tools, SummarizationMiddleware)
│ ├── Depth 2: Tool calls (sql_db_query, retriever, etc.)
│ └── Depth 2: LLM calls (ChatOpenAI, ChatAnthropic)
└── Depth 3+: Nested subagent calls
Use --root-run-name to target specific agent frameworks:
--root-run-name LangGraphFull conversation with expected output - tests complete agent behavior.
# Basic usage
python generate_datasets.py --type final_response \
--project my-project \
--root-run-name LangGraph \
--limit 30 \
--output /tmp/final_response.json
# With custom output fields
python generate_datasets.py --type final_response \
--project my-project \
--output-fields "answer,result" \
--output /tmp/final.json
# Messages only (ignore output dict keys)
python generate_datasets.py --type final_response \
--project my-project \
--messages-only \
--output /tmp/final.json
Structure:
{
"trace_id": "...",
"inputs": {"query": "What are the top 3 genres?"},
"outputs": {
"expected_response": "The top 3 genres based on the number of tracks are:\n\n1. Rock with 1,297 tracks\n2. Latin with 579 tracks\n3. Metal with 374 tracks"
}
}
Extraction Priority:
--output-fields)Important: Always checks root run first for final response to avoid intermediate tool outputs.
Single node inputs/outputs - tests any specific node's behavior. Supports multiple occurrences per trace to capture conversation evolution.
# Extract all occurrences (default)
python generate_datasets.py --type single_step \
--project my-project \
--root-run-name LangGraph \
--run-name model \
--output /tmp/single_step.json
# Sample 2 occurrences per trace
python generate_datasets.py --type single_step \
--project my-project \
--root-run-name LangGraph \
--run-name model \
--sample-per-trace 2 \
--output /tmp/single_step_sampled.json
# Target specific tool at depth 2
python generate_datasets.py --type single_step \
--project my-project \
--root-run-name LangGraph \
--run-name sql_db_query \
--output /tmp/sql_query.json
Structure:
{
"trace_id": "...",
"run_id": "...",
"occurrence": 2,
"inputs": {
"messages": [
{"type": "human", "content": "What are the top 3 genres?"},
{"type": "ai", "content": "", "tool_calls": [...]},
{"type": "tool", "content": "...results..."},
...
]
},
"outputs": {
"expected_output": {
"messages": [
{"type": "ai", "content": "", "tool_calls": [...]}
]
},
"node_name": "model"
}
}
Key Features:
occurrence field tracks which invocation (1st, 2nd, 3rd, etc.)--sample-per-trace randomly samples N occurrences per trace--run-name to target any node at any depthCommon targets:
model (depth 1) - LLM invocations with growing contexttools (depth 1) - Tool execution chainTool call sequence - tests execution path with configurable depth.
# Include all tool calls (all depths)
python generate_datasets.py --type trajectory \
--project my-project \
--root-run-name LangGraph \
--limit 30 \
--output /tmp/trajectory_all.json
# Only tool calls up to depth 2
python generate_datasets.py --type trajectory \
--project my-project \
--root-run-name LangGraph \
--depth 2 \
--output /tmp/trajectory_depth2.json
# Only root-level tool calls (depth 0) - usually empty if tools are at depth 2+
python generate_datasets.py --type trajectory \
--project my-project \
--depth 0 \
--output /tmp/trajectory_root.json
Structure:
{
"trace_id": "...",
"inputs": {"query": "What are the top 3 genres?"},
"outputs": {
"expected_trajectory": [
"sql_db_list_tables",
"sql_db_schema",
"sql_db_query_checker",
"sql_db_query"
]
}
}
Depth Control:
--depth = all levels (includes subagent tool calls)--depth 2 = root + 2 levels (typical for capturing all main tools)--depth 1 = often only middleware/chains, no actual tool calls--depth 0 = root only (no tool calls)Note: Tool calls are typically at depth 2 in LangGraph/DeepAgents architecture.
Question/chunks/answer/citations - tests retrieval quality.
python generate_datasets.py --type rag \
--project my-project \
--limit 30 \
--output /tmp/rag_ds.csv # Supports .json or .csv
Structure (CSV format):
question,retrieved_chunks,answer,cited_chunks
"How do I...","Chunk 1\n\nChunk 2","The answer is...","[\"Chunk 1\"]"
All dataset types support both JSON and CSV:
# JSON output (default)
python generate_datasets.py --type trajectory --project my-project --output ds.json
# CSV output (use .csv extension)
python generate_datasets.py --type trajectory --project my-project --output ds.csv
# Generate and upload in one command
python generate_datasets.py --type trajectory \
--project my-project \
--root-run-name LangGraph \
--limit 50 \
--output /tmp/trajectory_ds.json \
--upload "Skills: Trajectory"
# Use --replace to overwrite existing dataset
python generate_datasets.py --type final_response \
--project my-project \
--output /tmp/final.json \
--upload "Skills: Final Response" \
--replace
Naming Convention: Use "Skills: <Type>" format for consistency:
# List all datasets
python query_datasets.py list-datasets
# Filter by name pattern
python query_datasets.py list-datasets | grep "Skills:"
# View dataset examples
python query_datasets.py show "Skills: Trajectory" --limit 5
# View local file
python query_datasets.py view-file /tmp/trajectory_ds.json --limit 3
# Analyze structure
python query_datasets.py structure /tmp/trajectory_ds.json
# Export from LangSmith to local
python query_datasets.py export "Skills: Final Response" /tmp/exported.json --limit 100
--root-run-name - Filter for specific agent framework (e.g., "LangGraph")--last-n-minutes 1440 for last 24 hours of data--sample-per-trace 2 to capture conversation evolution--depth 2 typically captures all main tool callsquery_datasets.py view-file to inspect first--replace carefully - Overwrites existing datasets, useful for iteration# 1. Generate fresh traces (if needed)
python tests/test_agent.py --batch # Your test agent
# 2. Generate all dataset types from LangGraph traces
python generate_datasets.py --type final_response \
--project skills --root-run-name LangGraph --limit 10 \
--output /tmp/final.json --upload "Skills: Final Response" --replace
python generate_datasets.py --type single_step \
--project skills --root-run-name LangGraph --run-name model \
--sample-per-trace 2 --limit 10 \
--output /tmp/model.json --upload "Skills: Single Step (model)" --replace
python generate_datasets.py --type trajectory \
--project skills --root-run-name LangGraph --limit 10 \
--output /tmp/traj.json --upload "Skills: Trajectory (all depths)" --replace
python generate_datasets.py --type trajectory \
--project skills --root-run-name LangGraph --depth 2 --limit 10 \
--output /tmp/traj_d2.json --upload "Skills: Trajectory (depth=2)" --replace
# 3. Review in LangSmith UI
# Visit https://smith.langchain.com → Datasets → Filter for "Skills:"
# 4. Query locally if needed
python query_datasets.py show "Skills: Final Response" --limit 3
Empty final_response outputs:
--root-run-name matches your agent's root node--messages-only if output dict is emptyNo trajectory examples:
--depth or use --depth 2python query_traces.py trace <id> --show-hierarchyToo many single_step examples:
--sample-per-trace 2 to limit examples per traceDataset upload fails:
--replaceSearch for places (restaurants, cafes, etc.) via Google Places API proxy on localhost.
Interact with GitHub using the `gh` CLI. Use `gh issue`, `gh pr`, `gh run`, and `gh api` for issues, PRs, CI runs, and advanced queries.
Create or update AgentSkills. Use when designing, structuring, or packaging skills with scripts, references, and assets.
Start voice calls via the OpenClaw voice-call plugin.
Notion API for creating and managing pages, databases, and blocks.
Gemini CLI for one-shot Q&A, summaries, and generation.
Category:developer