๐ฌ Local LLM OptimizationJune 23, 2026โ
Tests passing
LLM Batch Inference Tool
A CLI utility to optimize and batch process input for faster inference on large language models. It intelligently splits large input datasets into manageable chunks, manages token limits, and parallelizes inference requests for local LLMs, significantly reducing processing time.
What It Does
- Supports text and JSON input files.
- Splits input data into manageable chunks based on a specified batch size.
- Performs parallelized inference using the Hugging Face
transformerslibrary. - Saves the output in JSON format, including timing information.
Installation
- Python 3.7+
transformerslibraryjobliblibrarypytestfor testing
Install the required dependencies using pip:
pip install transformers joblib pytestUsage
Run the tool from the command line:
python llm_batch_infer_tool.py --input <input_file> --output <output_file> [--batch_size <batch_size>] [--parallel <parallel_processes>]Arguments
--input: Path to the input file (text or JSON).--output: Path to save the output JSON file.--batch_size: (Optional) Batch size for processing. Default is 16.--parallel: (Optional) Number of parallel processes. Default is 4.
Example
python llm_batch_infer_tool.py --input data/input.txt --output data/output.json --batch_size 8 --parallel 2Source Code
import argparse
import json
import time
from pathlib import Path
from typing import List, Union
from joblib import Parallel, delayed
from transformers import pipeline
# Constants
DEFAULT_BATCH_SIZE = 16
DEFAULT_PARALLEL = 4
MODEL_NAME = "gpt2"
TOKEN_LIMIT = 512
def load_input(file_path: str) -> Union[List[str], List[dict]]:
"""Load input data from a text or JSON file."""
path = Path(file_path)
if not path.exists():
raise FileNotFoundError(f"Input file {file_path} does not exist.")
if path.suffix == ".txt":
with open(file_path, "r", encoding="utf-8") as f:
return [line.strip() for line in f if line.strip()]
elif path.suffix == ".json":
with open(file_path, "r", encoding="utf-8") as f:
return json.load(f)
else:
raise ValueError("Unsupported file format. Only .txt and .json are supported.")
def save_output(output: Union[List[str], List[dict]], output_path: str):
"""Save output data to a JSON file."""
with open(output_path, "w", encoding="utf-8") as f:
json.dump(output, f, indent=4)
def chunk_input(data: List[str], batch_size: int) -> List[List[str]]:
"""Split input data into chunks of specified batch size."""
return [data[i:i + batch_size] for i in range(0, len(data), batch_size)]
def process_batch(batch: List[str], model_pipeline) -> List[str]:
"""Process a single batch of inputs using the LLM pipeline."""
return [result['generated_text'] for result in model_pipeline(batch, max_length=TOKEN_LIMIT)]
def batch_inference(input_data: List[str], batch_size: int, parallel: int) -> List[str]:
"""Perform batch inference using the specified batch size and parallelization."""
model_pipeline = pipeline("text-generation", model=MODEL_NAME, tokenizer=MODEL_NAME, device=-1)
chunks = chunk_input(input_data, batch_size)
results = Parallel(n_jobs=parallel)(
delayed(process_batch)(chunk, model_pipeline) for chunk in chunks
)
return [item for sublist in results for item in sublist]
def main():
parser = argparse.ArgumentParser(
description="LLM Batch Inference Tool: Process input data in batches for faster inference using large language models."
)
parser.add_argument("--input", required=True, help="Path to the input file (text or JSON).")
parser.add_argument("--output", required=True, help="Path to save the output JSON file.")
parser.add_argument("--batch_size", type=int, default=DEFAULT_BATCH_SIZE, help="Batch size for processing (default: 16).")
parser.add_argument("--parallel", type=int, default=DEFAULT_PARALLEL, help="Number of parallel processes (default: 4).")
args = parser.parse_args()
try:
input_data = load_input(args.input)
if not input_data:
raise ValueError("Input file is empty.")
start_time = time.time()
results = batch_inference(input_data, args.batch_size, args.parallel)
end_time = time.time()
output_data = {
"results": results,
"timing": {
"total_time_seconds": end_time - start_time,
"batch_size": args.batch_size,
"parallel": args.parallel
}
}
save_output(output_data, args.output)
print(f"Processing completed. Results saved to {args.output}")
except Exception as e:
print(f"Error: {e}")
if __name__ == "__main__":
main()
Community
Downloads
ยทยทยท
Rate this tool
No ratings yet โ be the first!
Details
- Tool Name
- llm_batch_infer_tool
- Category
- Local LLM Optimization
- Generated
- June 23, 2026
- Tests
- Passing โ
- Fix Loops
- 3
Quick Install
Clone just this tool:
git clone --depth 1 --filter=blob:none --sparse \ https://github.com/ptulin/autoaiforge.git cd autoaiforge git sparse-checkout set generated_tools/2026-06-23/llm_batch_infer_tool cd generated_tools/2026-06-23/llm_batch_infer_tool pip install -r requirements.txt 2>/dev/null || true python llm_batch_infer_tool.py