All Toolsโ€บReal-Time Inference Profiler
๐Ÿ”ง Low-Latency AI InferenceJuly 29, 2026โœ… Tests passing

Real-Time Inference Profiler

This tool helps developers simulate and profile their AI models under real-time conditions by generating load scenarios (e.g., changing request rates) and measuring performance metrics such as latency, throughput, and memory usage. It provides insights into how models behave under different scalability demands.

What It Does

  • Simulate steady and bursty workloads.
  • Measure performance metrics like latency and throughput.
  • Supports multiple concurrent workers.
  • Outputs results in JSON or CSV format.

Installation

Install the required dependencies:

pip install numpy pandas pytest

Usage

Command-Line Interface

Run the profiler from the command line:

python realtime_inference_profiler.py --workload steady --rate 10 --duration 10 --workers 2 --output-format json

Programmatic Usage

You can also use the profiler programmatically:

from realtime_inference_profiler import profile

def my_model():
    # Simulate model inference
    time.sleep(0.01)

metrics = profile(my_model, workload='steady', rate=10, duration=10, workers=2, output_format='json')
print(metrics)

Source Code

import asyncio
import time
import numpy as np
import pandas as pd
from typing import Callable, Dict, Any

def generate_workload(workload: str, rate: int, duration: int) -> asyncio.Queue:
    """
    Generate a workload queue based on the specified pattern.

    Args:
        workload (str): The workload pattern ('steady' or 'bursty').
        rate (int): Requests per second.
        duration (int): Duration of the workload in seconds.

    Returns:
        asyncio.Queue: A queue with timestamps for requests.
    """
    queue = asyncio.Queue()
    if workload == 'steady':
        interval = 1 / rate
        for i in range(rate * duration):
            queue.put_nowait(time.time() + i * interval)
    elif workload == 'bursty':
        for i in range(duration):
            burst_size = np.random.poisson(rate)
            for _ in range(burst_size):
                queue.put_nowait(time.time() + i)
    else:
        raise ValueError("Unsupported workload type. Use 'steady' or 'bursty'.")
    return queue

async def worker(queue: asyncio.Queue, model: Callable, results: list):
    """
    Process requests from the queue and record performance metrics.

    Args:
        queue (asyncio.Queue): The workload queue.
        model (Callable): The inference function to profile.
        results (list): A list to store the results.
    """
    while not queue.empty():
        request_time = await queue.get()
        await asyncio.sleep(max(0, request_time - time.time()))
        start_time = time.time()
        try:
            model()  # Call the model function
        except Exception as e:
            results.append({'error': str(e)})
            continue
        end_time = time.time()
        results.append({'latency': end_time - start_time, 'timestamp': start_time})

async def profile_async(model: Callable, workload: str, rate: int, duration: int, workers: int) -> pd.DataFrame:
    """
    Profile the performance of a model under a specific workload.

    Args:
        model (Callable): The inference function to profile.
        workload (str): The workload pattern ('steady' or 'bursty').
        rate (int): Requests per second.
        duration (int): Duration of the workload in seconds.
        workers (int): Number of concurrent workers.

    Returns:
        pd.DataFrame: Performance metrics.
    """
    queue = generate_workload(workload, rate, duration)
    results = []
    tasks = [worker(queue, model, results) for _ in range(workers)]
    await asyncio.gather(*tasks)
    return pd.DataFrame(results)

def profile(model: Callable, workload: str, rate: int, duration: int = 10, workers: int = 1, output_format: str = 'json') -> Any:
    """
    Profile the performance of a model and return metrics.

    Args:
        model (Callable): The inference function to profile.
        workload (str): The workload pattern ('steady' or 'bursty').
        rate (int): Requests per second.
        duration (int): Duration of the workload in seconds.
        workers (int): Number of concurrent workers.
        output_format (str): Output format ('json' or 'csv').

    Returns:
        Any: Performance metrics in the specified format.
    """
    loop = asyncio.new_event_loop()
    asyncio.set_event_loop(loop)
    metrics = loop.run_until_complete(profile_async(model, workload, rate, duration, workers))
    loop.close()
    if output_format == 'json':
        return metrics.to_dict(orient='records')
    elif output_format == 'csv':
        return metrics.to_csv(index=False)
    else:
        raise ValueError("Unsupported output format. Use 'json' or 'csv'.")

if __name__ == "__main__":
    import argparse

    def dummy_model():
        """A dummy model for testing."""
        time.sleep(0.01)

    parser = argparse.ArgumentParser(description="Real-Time Inference Profiler")
    parser.add_argument('--workload', type=str, required=True, choices=['steady', 'bursty'], help="Workload pattern")
    parser.add_argument('--rate', type=int, required=True, help="Requests per second")
    parser.add_argument('--duration', type=int, default=10, help="Duration of the workload in seconds")
    parser.add_argument('--workers', type=int, default=1, help="Number of concurrent workers")
    parser.add_argument('--output-format', type=str, default='json', choices=['json', 'csv'], help="Output format")

    args = parser.parse_args()

    metrics = profile(dummy_model, args.workload, args.rate, args.duration, args.workers, args.output_format)
    if args.output_format == 'json':
        print(metrics)
    else:
        print(metrics)

Community

Downloads

ยทยทยท

Rate this tool

No ratings yet โ€” be the first!

Details

Tool Name
realtime_inference_profiler
Category
Low-Latency AI Inference
Generated
July 29, 2026
Tests
Passing โœ…
Fix Loops
3

Quick Install

Clone just this tool:

git clone --depth 1 --filter=blob:none --sparse \
  https://github.com/ptulin/autoaiforge.git
cd autoaiforge
git sparse-checkout set generated_tools/2026-07-29/realtime_inference_profiler
cd generated_tools/2026-07-29/realtime_inference_profiler
pip install -r requirements.txt 2>/dev/null || true
python realtime_inference_profiler.py
Real-Time Inference Profiler โ€” AI Tools by AutoAIForge