All Toolsโ€บContext Window Evaluator
๐Ÿ’ฌ LLM Context Optimization ToolsJuly 12, 2026โœ… Tests passing

Context Window Evaluator

This tool analyzes prompt datasets to evaluate and optimize how effectively they utilize the context window of large language models (LLMs). It provides insights on token density, redundancy, and truncation risks, helping developers maximize information fit within the LLM's limits.

What It Does

  • Analyze token usage, redundancy, and truncation risks for individual prompts.
  • Process multiple prompts from text or JSON files.
  • Generate a detailed report in the terminal or save it as a JSON file.

Installation

Install the required dependencies using pip:

pip install tiktoken nltk rich

Usage

Command Line Interface

Run the tool from the command line:

python context_window_evaluator.py --input <file1> <file2> --output <output_file>
  • --input: One or more input files (text or JSON) containing prompts to analyze.
  • --output: (Optional) Path to save the analysis report as a JSON file.

Example

Analyze a file prompts.txt and save the report to report.json:

python context_window_evaluator.py --input prompts.txt --output report.json

Source Code

import argparse
import json
from pathlib import Path
from rich.console import Console
from rich.table import Table
import tiktoken
import nltk
from nltk.tokenize import word_tokenize

def analyze_prompt(prompt, tokenizer):
    """
    Analyze a single prompt for token usage, redundancy, and truncation risks.

    Args:
        prompt (str): The text of the prompt to analyze.
        tokenizer: The tokenizer instance from tiktoken.

    Returns:
        dict: Analysis results including token count, unique token count, and redundancy.
    """
    tokens = tokenizer.encode(prompt)
    token_count = len(tokens)
    unique_tokens = len(set(tokens))
    redundancy = 1 - (unique_tokens / token_count) if token_count > 0 else 0

    return {
        "token_count": token_count,
        "unique_tokens": unique_tokens,
        "redundancy": redundancy,
        "truncation_risk": token_count > tokenizer.max_token_count
    }

def analyze_file(file_path, tokenizer):
    """
    Analyze a file containing prompts.

    Args:
        file_path (str or Path): Path to the input file.
        tokenizer: The tokenizer instance from tiktoken.

    Returns:
        list: A list of analysis results for each prompt.
    """
    results = []
    try:
        file_path = Path(file_path)  # Ensure file_path is a Path object
        with open(file_path, 'r', encoding='utf-8') as file:
            if file_path.suffix == '.json':
                data = json.load(file)
                prompts = data if isinstance(data, list) else [data]
            else:
                prompts = file.readlines()

            for prompt in prompts:
                prompt = prompt.strip()
                if prompt:
                    results.append(analyze_prompt(prompt, tokenizer))
    except Exception as e:
        raise ValueError(f"Error processing file {file_path}: {e}")

    return results

def generate_report(analysis_results, output_file=None):
    """
    Generate a report from the analysis results.

    Args:
        analysis_results (list): List of analysis results.
        output_file (str, optional): Path to save the report as JSON.

    Returns:
        None
    """
    console = Console()
    table = Table(title="Context Window Analysis Report")

    table.add_column("Prompt #", justify="right", style="cyan")
    table.add_column("Token Count", justify="right", style="green")
    table.add_column("Unique Tokens", justify="right", style="magenta")
    table.add_column("Redundancy", justify="right", style="yellow")
    table.add_column("Truncation Risk", justify="right", style="red")

    for idx, result in enumerate(analysis_results, start=1):
        table.add_row(
            str(idx),
            str(result["token_count"]),
            str(result["unique_tokens"]),
            f"{result['redundancy']:.2f}",
            "Yes" if result["truncation_risk"] else "No"
        )

    console.print(table)

    if output_file:
        with open(output_file, 'w', encoding='utf-8') as f:
            json.dump(analysis_results, f, indent=4)

def main():
    parser = argparse.ArgumentParser(description="Context Window Evaluator")
    parser.add_argument('--input', required=True, nargs='+', help="Input text or JSON file(s) containing prompts.")
    parser.add_argument('--output', help="Optional output file to save the analysis report in JSON format.")
    args = parser.parse_args()

    nltk.download('punkt', quiet=True)
    tokenizer = tiktoken.get_encoding("cl100k_base")
    tokenizer.max_token_count = 4096

    all_results = []

    for input_file in args.input:
        file_path = Path(input_file)
        if not file_path.exists():
            print(f"Error: File not found - {file_path}")
            continue

        try:
            results = analyze_file(file_path, tokenizer)
            all_results.extend(results)
        except ValueError as e:
            print(e)

    generate_report(all_results, args.output)

if __name__ == "__main__":
    main()

Community

Downloads

ยทยทยท

Rate this tool

No ratings yet โ€” be the first!

Details

Tool Name
context_window_evaluator
Category
LLM Context Optimization Tools
Generated
July 12, 2026
Tests
Passing โœ…
Fix Loops
2

Quick Install

Clone just this tool:

git clone --depth 1 --filter=blob:none --sparse \
  https://github.com/ptulin/autoaiforge.git
cd autoaiforge
git sparse-checkout set generated_tools/2026-07-12/context_window_evaluator
cd generated_tools/2026-07-12/context_window_evaluator
pip install -r requirements.txt 2>/dev/null || true
python context_window_evaluator.py
Context Window Evaluator โ€” AI Tools by AutoAIForge