{"uid":"cap_Zjr8p3W7uBetzEH_0PD5K","slug":"llm-eval-case-normalizer-4e2fae40","name":"LLM Eval Case Normalizer","description":"Explore 330 pay-per-call x402 API services and 27 agent-native digital products, with Base USDC pricing, secure Polar checkout, and free discovery.","url":"https://signalharness.ai/api/agent/services/llm_eval_case_normalize/invoke","method":"POST","headers":{},"bodySchema":{"type":"object","properties":{"request_json":{"type":"string","maxLength":65536,"minLength":2}}},"responseSchema":{"type":"json","example":{"replay":false,"result":{"warnings":["Verify the caller-supplied data before relying on this result."],"service_id":"llm_eval_case_normalize","analysis_json":"{\"example\":\"schema-valid caller-supplied data\"}","evidence_scope":"caller_supplied_data"},"status":"succeeded","receipt":{"asset":"0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913","usage":[],"status":"succeeded","network":"eip155:8453","artifacts":[],"endedAtMs":0,"latencyMs":0,"paymentId":"example-payment","receiptId":"example-receipt","requestId":"example-request","serviceId":"llm_eval_case_normalize","executionId":"example-execution","startedAtMs":0,"amountAtomic":"5000","resultSha256":"35c7edda0781047359e02190ab429a4a8847abb8649cc5048ab43265ce0ebcbe","serviceVersion":"1.0.0","settlementReference":"0x0000000000000000000000000000000000000000000000000000000000000000"},"artifacts":[],"requestId":"example-request","executionId":"example-execution"}},"example":null,"exampleRequest":null,"tags":["x402"],"displayCostAmount":"0.005","displayCostAsset":"USDC","priceDynamic":false,"priceHint":null,"priceStatus":"priced","priceSource":"probe","requiresHandshake":false,"reviewCount":0,"rating":{"score":"0.00","successRate":"0.00","reviews":0,"stars":null,"state":"unrated"},"availabilityStatus":"unknown","priceObserved":null,"sessionDeposit":null,"pricing":{"kind":"static","summary":"$0.005/call","primary":{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.005","per":"call","confidence":"exact"},"accepted":[{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.005","per":"call","confidence":"exact"}]},"paymentMethods":[{"uid":"pm_7zvW1bYfO5Yw4pvmcPGWS","protocol":"x402","methodType":"crypto","chain":"base","mode":"charge","costAmount":"0.005","costPer":"request","priority":0,"asset":"0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913","unit":"request","depositMicros":null,"planRef":null}],"brandName":null,"brandSlug":null,"brandBaseUrl":null,"brandDocsUrl":null,"whatItDoes":"Normalizes and standardizes LLM evaluation test case JSON into a canonical format for consistent evaluation pipelines.","exampleAgentPrompt":"Can you normalize this LLM evaluation case JSON for me so it's in a consistent, canonical format ready for benchmarking? Here's the raw JSON: {\"prompt\": \"What is 2+2?\", \"expected\": \"4\", \"tags\": [\"math\"]}","exampleUseCases":[{"title":"Standardizing eval dataset before benchmarking","prompt":"I have a bunch of LLM test cases in different formats — can you normalize this one into a canonical structure so it works with our eval pipeline? Here's the JSON: {\"input\": \"Summarize this article\", \"ground_truth\": \"Article summary\", \"category\": \"summarization\"}"},{"title":"Cleaning eval cases from a third-party dataset","prompt":"We got evaluation cases from an external vendor and they're not in our standard format. Can you run this eval case JSON through the normalizer to get it into canonical form? The JSON is: {\"q\": \"Who wrote Hamlet?\", \"a\": \"Shakespeare\", \"difficulty\": \"easy\"}"},{"title":"Validating eval case structure before pipeline ingestion","prompt":"Before I ingest these LLM evaluation cases into our benchmark pipeline, can you normalize and validate this JSON to make sure it's properly structured? Here it is: {\"user_message\": \"Translate to French: Hello\", \"expected_output\": \"Bonjour\", \"lang\": \"fr\"}"}],"resultDescription":"Returns a JSON object containing a normalized analysis_json string representing the standardized evaluation case, along with any warnings about the input data, the service_id, and an evidence_scope field indicating the data provenance. Also includes a payment receipt with execution metadata.","failureModes":["Invalid or malformed JSON string in request_json — returns error or warnings","Input JSON exceeds 65536 character limit — request rejected","Empty or too-short input JSON (minLength 2) — request rejected","Caller-supplied data that cannot be interpreted as a valid eval case — may produce warnings in the result"],"whenToPreferThis":"Use this endpoint when you need to normalize LLM evaluation test cases into a canonical format before feeding them into an evaluation or benchmarking pipeline. Prefer this over manual preprocessing when dealing with heterogeneous eval case formats from multiple sources or vendors. It is particularly useful in automated agent workflows that assemble eval datasets from diverse inputs.","instructions":null,"reviewSummary":null,"reviewSummaryHighlights":null,"reviewSummaryConcerns":null,"reviewSummaryGeneratedAt":null,"activationCount":0,"lastUsedAt":null,"lastSuccessfullyRanAt":null,"lastHealthCheckAt":"2026-09-13T07:02:42.781Z","isFirstParty":false}