affaan-m/ECC

regex-vs-llm-structured-text

構造化テキストの解析に正規表現と大規模言語モデルのどちらを使うかを選択するための意思決定フレームワーク——まず正規表達式から始め、信頼度の低いエッジケースにのみ大規模言語モデルを追加する。

68Collecting
See how to use itView GitHub source
npx skills add https://github.com/affaan-m/ECC --skill "docs/ja-JP/skills/regex-vs-llm-structured-text"
Automated source guide

Source checked Jul 28, 2026·Refresh due Oct 26, 2026

Reorganized from the pinned upstream SKILL.md

Turn regex-vs-llm-structured-text's source instructions into a guide you can follow

According to the pinned SKILL.md from affaan-m/ECC: 構造化テキスト(クイズ、フォーム、請求書、ドキュメント)を解析するための実用的な意思決定フレームワーク。核心的な洞察:正規表現は低コストかつ決定論的に95〜98%のケースを処理できる。コストのかかるLLM呼び出しは残りのエッジケースに留める。

npx skills add https://github.com/affaan-m/ECC --skill "docs/ja-JP/skills/regex-vs-llm-structured-text"
Check the pinned source

Best fit

  • 構造化テキストの解析に正規表現と大規模言語モデルのどちらを使うかを選択するための意思決定フレームワーク——まず正規表達式から始め、信頼度の低いエッジケースにのみ大規模言語モデルを追加する。

Bring this context

  • A concrete task that matches the documented purpose of regex-vs-llm-structured-text.
  • The files, examples, or context the task depends on.
  • Your constraints, target environment, and definition of done.

Expected outputs

  • A result that follows the pinned regex-vs-llm-structured-text instructions.
  • A concise record of assumptions, inputs used, and unresolved questions.
  • A final check against the source workflow and relevant permission signals.

Key source sections

Read regex-vs-llm-structured-text through these 5 source sections

Sections are extracted automatically from the pinned SKILL.md and link back to the source.

01

使用場面

繰り返しパターンを持つ構造化テキスト(設問、フォーム、表)の解析 テキスト抽出に正規表現とLLMのどちらを使うかの判断 両方のアプローチを組み合わせたハイブリッドパイプラインの構築 テキスト処理におけるコスト/精度のトレードオフの最適化

SKILL.md · 使用場面
繰り返しパターンを持つ構造化テキスト(設問、フォーム、表)の解析テキスト抽出に正規表現とLLMのどちらを使うかの判断両方のアプローチを組み合わせたハイブリッドパイプラインの構築
02

意思決定フレームワーク

Review the “意思決定フレームワーク” section in the pinned source before continuing.

SKILL.md · 意思決定フレームワーク
Review and apply the “意思決定フレームワーク” source section.
03

アーキテクチャパターン

Review the “アーキテクチャパターン” section in the pinned source before continuing.

SKILL.md · アーキテクチャパターン
Review and apply the “アーキテクチャパターン” source section.
04

実装

LLMによるレビューが必要かもしれない項目にフラグを立てる:

SKILL.md · 実装
LLMによるレビューが必要かもしれない項目にフラグを立てる:
05

1. 正規表現パーサー(大半のケースを処理)

Review the “1. 正規表現パーサー(大半のケースを処理)” section in the pinned source before continuing.

SKILL.md · 1. 正規表現パーサー(大半のケースを処理)
Review and apply the “1. 正規表現パーサー(大半のケースを処理)” source section.

SkillSignal prompt templates

Provide the task, context, and acceptance criteria

These prompts were written by SkillSignal from the source structure; they are not upstream text.

Task-start prompt

Confirm source fit, inputs, and outputs before acting.

Use regex-vs-llm-structured-text to help me with: [specific task]. Context: [files, data, or background]. Constraints: [environment, scope, and prohibited actions]. Before acting, check the pinned SKILL.md and explain which sections apply, what inputs are still missing, and what you will deliver.

Source-guided execution

Make the Agent explicitly follow the key extracted sections.

Apply the pinned regex-vs-llm-structured-text source to [task]. Pay particular attention to these source sections: “使用場面”, “意思決定フレームワーク”, “アーキテクチャパターン”, “実装”, “1. 正規表現パーサー(大半のケースを処理)”. Preserve the important decision at each step. Mark facts not covered by the source as “needs confirmation” instead of inventing them. Then verify the result against my acceptance criteria: [criteria].

Result-review prompt

Check omissions, permissions, and source drift before delivery.

Review the current regex-vs-llm-structured-text result: (1) does it satisfy the original task; (2) were any applicable steps or limits in the pinned SKILL.md missed; (3) did it perform any unauthorized file, command, network, or data action; and (4) which conclusions remain unverified? List issues first, then fix only what the source or user authorization supports.

Output checklist

Verify each item before delivery

The task matches the purpose documented in the SKILL.md.

The source section “使用場面” has been checked.

The source section “意思決定フレームワーク” has been checked.

The source section “アーキテクチャパターン” has been checked.

The source section “実装” has been checked.

Inputs, constraints, and acceptance criteria are explicit.

Unverified facts, compatibility, and outcome claims are clearly marked.

Any file, command, network, or data action has been reviewed.

Choose a different workflow

When another Skill is the better fit

FAQ

What does regex-vs-llm-structured-text do?

構造化テキスト(クイズ、フォーム、請求書、ドキュメント)を解析するための実用的な意思決定フレームワーク。核心的な洞察:正規表現は低コストかつ決定論的に95〜98%のケースを処理できる。コストのかかるLLM呼び出しは残りのエッジケースに留める。

How do I start using regex-vs-llm-structured-text?

The catalog detected this source-specific install command: npx skills add https://github.com/affaan-m/ECC --skill "docs/ja-JP/skills/regex-vs-llm-structured-text". Inspect the command and pinned source before running it.

Which Agent platforms does it declare?

No dedicated Agent platform is declared in the pinned source record.

Repository stars
234,327
Repository forks
35,711
Quality
68/100
Source repository last pushed

Quality breakdown

Based on traceable docs and repository signals; stars are not treated as quality.

68/100
Documentation25/30
Specificity11/25
Maintenance20/20
Trust signals12/25

Compare before choosing

Related Agent Skills and source variants

These links are selected from shared tasks, functions, stacks, platforms, and same-name variants. Compare the source owner, documentation, permissions, and maintenance signals.

regex-vs-llm-structured-text by affaan-m

Decision framework for choosing between regex and LLM when parsing structured text — start with regex, add LLM only for low-confidence edge cases.

regex-vs-llm-structured-text by affaan-m

选择在解析结构化文本时使用正则表达式还是大型语言模型的决策框架——从正则表达式开始,仅在低置信度的边缘情况下添加大型语言模型。

ab-testing by coreyhaines31

When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program

churn-prevention by coreyhaines31

When the user wants to reduce churn, build cancellation flows, set up save offers, recover failed payments, or implement retention strategies. Also use when the user mentions 'churn,' 'cancel flow,' 'offboarding,' 'save offer,' 'dunning,' 'failed payment recovery,' 'win-back,' 'retention,' 'exit survey,' 'pause subscription,' 'involuntary churn,' 'people keep canceling,' 'churn rate is too high,' 'how do I keep users,' or 'customers are leaving.' Use this whenever someone is losing subscribers o

design-intelligence by event4u-app

Grounded design brief from the adopted corpus — style, WCAG-checked color tokens, typography, layout pattern, anti-patterns. Use on ui-design-brief or any which-style/palette/font/chart decision.

View original Skill.mdThis page is parsed directly from the repository SKILL.md without editorial rewriting. Collected: Jul 28, 2026 · about 1 min

構造化テキスト解析における正規表現 vs LLM

構造化テキスト(クイズ、フォーム、請求書、ドキュメント)を解析するための実用的な意思決定フレームワーク。核心的な洞察:正規表現は低コストかつ決定論的に95〜98%のケースを処理できる。コストのかかるLLM呼び出しは残りのエッジケースに留める。

使用場面

  • 繰り返しパターンを持つ構造化テキスト(設問、フォーム、表)の解析
  • テキスト抽出に正規表現とLLMのどちらを使うかの判断
  • 両方のアプローチを組み合わせたハイブリッドパイプラインの構築
  • テキスト処理におけるコスト/精度のトレードオフの最適化

意思決定フレームワーク

テキスト形式は一貫していて繰り返しがあるか?
├── はい (>90% が何らかのパターンに従う) → 正規表現から始める
│   ├── 正規表現が 95%+ を処理 → 完了、LLM は不要
│   └── 正規表現が <95% を処理 → エッジケースのみ LLM を追加
└── いいえ (自由形式、高度に可変) → LLM を直接使用

アーキテクチャパターン

[正規表現パーサー] ─── 構造を抽出(95〜98% の精度)
    │
    ▼
[テキストクリーナー] ─── ノイズを除去(マーカー、ページ番号、アーティファクト)
    │
    ▼
[信頼度スコアラー] ─── 信頼度の低い抽出結果にフラグを立てる
    │
    ├── 高信頼度(≥0.95)→ 直接出力
    │
    └── 低信頼度(<0.95)→ [LLM バリデーター] → 出力

実装

1. 正規表現パーサー(大半のケースを処理)

import re
from dataclasses import dataclass

@dataclass(frozen=True)
class ParsedItem:
    id: str
    text: str
    choices: tuple[str, ...]
    answer: str
    confidence: float = 1.0

def parse_structured_text(content: str) -> list[ParsedItem]:
    """Parse structured text using regex patterns."""
    pattern = re.compile(
        r"(?P<id>\d+)\.\s*(?P<text>.+?)\n"
        r"(?P<choices>(?:[A-D]\..+?\n)+)"
        r"Answer:\s*(?P<answer>[A-D])",
        re.MULTILINE | re.DOTALL,
    )
    items = []
    for match in pattern.finditer(content):
        choices = tuple(
            c.strip() for c in re.findall(r"[A-D]\.\s*(.+)", match.group("choices"))
        )
        items.append(ParsedItem(
            id=match.group("id"),
            text=match.group("text").strip(),
            choices=choices,
            answer=match.group("answer"),
        ))
    return items

2. 信頼度スコアリング

LLMによるレビューが必要かもしれない項目にフラグを立てる:

@dataclass(frozen=True)
class ConfidenceFlag:
    item_id: str
    score: float
    reasons: tuple[str, ...]

def score_confidence(item: ParsedItem) -> ConfidenceFlag:
    """Score extraction confidence and flag issues."""
    reasons = []
    score = 1.0

    if len(item.choices) < 3:
        reasons.append("few_choices")
        score -= 0.3

    if not item.answer:
        reasons.append("missing_answer")
        score -= 0.5

    if len(item.text) < 10:
        reasons.append("short_text")
        score -= 0.2

    return ConfidenceFlag(
        item_id=item.id,
        score=max(0.0, score),
        reasons=tuple(reasons),
    )

def identify_low_confidence(
    items: list[ParsedItem],
    threshold: float = 0.95,
) -> list[ConfidenceFlag]:
    """Return items below confidence threshold."""
    flags = [score_confidence(item) for item in items]
    return [f for f in flags if f.score < threshold]

3. LLM バリデーター(エッジケースのみ)

def validate_with_llm(
    item: ParsedItem,
    original_text: str,
    client,
) -> ParsedItem:
    """Use LLM to fix low-confidence extractions."""
    response = client.messages.create(
        model="claude-haiku-4-5-20251001",  # Cheapest model for validation
        max_tokens=500,
        messages=[{
            "role": "user",
            "content": (
                f"Extract the question, choices, and answer from this text.\n\n"
                f"Text: {original_text}\n\n"
                f"Current extraction: {item}\n\n"
                f"Return corrected JSON if needed, or 'CORRECT' if accurate."
            ),
        }],
    )
    # Parse LLM response and return corrected item...
    return corrected_item

4. ハイブリッドパイプライン

def process_document(
    content: str,
    *,
    llm_client=None,
    confidence_threshold: float = 0.95,
) -> list[ParsedItem]:
    """Full pipeline: regex -> confidence check -> LLM for edge cases."""
    # Step 1: Regex extraction (handles 95-98%)
    items = parse_structured_text(content)

    # Step 2: Confidence scoring
    low_confidence = identify_low_confidence(items, confidence_threshold)

    if not low_confidence or llm_client is None:
        return items

    # Step 3: LLM validation (only for flagged items)
    low_conf_ids = {f.item_id for f in low_confidence}
    result = []
    for item in items:
        if item.id in low_conf_ids:
            result.append(validate_with_llm(item, content, llm_client))
        else:
            result.append(item)

    return result

実際のメトリクス

本番のクイズ解析パイプライン(410項目)より:

メトリクス
正規表現の成功率98.0%
低信頼度項目8 (2.0%)
必要なLLM呼び出し回数~5
全件LLM比のコスト節約~95%
テストカバレッジ93%

ベストプラクティス

  • 正規表現から始める — 不完全な正規表現でも改善のベースラインになる
  • 信頼度スコアリングを使用して、LLMの助けが必要なものをプログラムで特定する
  • 最も安価なLLMを使用して検証する(Haikuクラスのモデルで十分)
  • 解析済み項目を変更しない — クリーニング/検証ステップから新しいインスタンスを返す
  • TDDは解析器に効果的 — まず既知のパターンのテストを書き、次にエッジケースを書く
  • メトリクスを記録(正規表現の成功率、LLM呼び出し回数)してパイプラインの健全性を追跡する

避けるべきアンチパターン

  • 正規表現が95%以上を処理できる場合に全テキストをLLMに送る(コスト高・低速)
  • 自由形式で高度に可変なテキストに正規表現を使用する(LLMの方が適切)
  • 信頼度スコアリングをスキップして正規表現が「うまくいく」ことを期待する
  • クリーニング/検証ステップで解析済みオブジェクトを変更する
  • エッジケースをテストしない(不正な入力、欠損フィールド、エンコーディング問題)

適用場面

  • クイズ/試験問題の解析
  • フォームデータの抽出
  • 請求書/レシートの処理
  • ドキュメント構造の解析(見出し、セクション、表)
  • 繰り返しパターンがあり、コストが重要なあらゆる構造化テキスト
Source repo
affaan-m/ECC
Skill path
docs/ja-JP/skills/regex-vs-llm-structured-text/SKILL.md
Commit SHA
4e973d3eaf92
Repository license
MIT
Data collected