{"uid":"cap_7ZXm7D_xMfHEpqv-ZwnEJ","slug":"signalharness-llm-eval-result-aggregator-18de0067","name":"SignalHarness LLM Eval Result Aggregator","description":"Explore 330 pay-per-call x402 API services and 27 agent-native digital products, with Base USDC pricing, secure Polar checkout, and free discovery.","url":"https://signalharness.ai/api/agent/services/llm_eval_result_aggregate/invoke","method":"POST","headers":{},"bodySchema":{"type":"object","properties":{"request_json":{"type":"string","maxLength":65536,"minLength":2}}},"responseSchema":{"type":"json","example":{"replay":false,"result":{"warnings":["Verify the caller-supplied data before relying on this result."],"service_id":"llm_eval_result_aggregate","analysis_json":"{\"example\":\"schema-valid caller-supplied data\"}","evidence_scope":"caller_supplied_data"},"status":"succeeded","receipt":{"asset":"0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913","usage":[],"status":"succeeded","network":"eip155:8453","artifacts":[],"endedAtMs":0,"latencyMs":0,"paymentId":"example-payment","receiptId":"example-receipt","requestId":"example-request","serviceId":"llm_eval_result_aggregate","executionId":"example-execution","startedAtMs":0,"amountAtomic":"10000","resultSha256":"5696e63c3b286d34c9110108be3b9a1437ddca4931ef8f7b1551b38e5e427c78","serviceVersion":"1.0.0","settlementReference":"0x0000000000000000000000000000000000000000000000000000000000000000"},"artifacts":[],"requestId":"example-request","executionId":"example-execution"}},"example":null,"exampleRequest":null,"tags":["x402"],"displayCostAmount":"0.01","displayCostAsset":"USDC","priceDynamic":false,"priceHint":null,"priceStatus":"priced","priceSource":"probe","requiresHandshake":false,"reviewCount":0,"rating":{"score":"0.00","successRate":"0.00","reviews":0,"stars":null,"state":"unrated"},"availabilityStatus":"unknown","priceObserved":null,"sessionDeposit":null,"pricing":{"kind":"static","summary":"$0.01/call","primary":{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.01","per":"call","confidence":"exact"},"accepted":[{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.01","per":"call","confidence":"exact"}]},"paymentMethods":[{"uid":"pm_-TgE46hHg84qi-QsvvJqx","protocol":"x402","methodType":"crypto","chain":"base","mode":"charge","costAmount":"0.01","costPer":"request","priority":0,"asset":"0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913","unit":"request","depositMicros":null,"planRef":null}],"brandName":null,"brandSlug":null,"brandBaseUrl":null,"brandDocsUrl":null,"whatItDoes":"Aggregates and summarizes evaluation results from LLM outputs, returning structured analysis JSON with warnings and evidence scope metadata.","exampleAgentPrompt":"Take this JSON blob of LLM evaluation results and aggregate them into a structured summary for me — I need the combined analysis with any warnings about data quality flagged: {\"model\":\"gpt-4\",\"scores\":[0.82,0.91,0.76],\"task\":\"summarization\"}.","exampleUseCases":[{"title":"Batch LLM benchmark rollup","prompt":"I ran evaluations on three different LLM responses for the same prompt and got separate scores — can you aggregate this JSON into a single consolidated result with any data quality warnings? Here's the eval data: {\"runs\":[{\"score\":0.88,\"model\":\"claude-3\"},{\"score\":0.75,\"model\":\"gpt-4\"},{\"score\":0.91,\"model\":\"gemini\"}]}."},{"title":"Model quality monitoring pipeline","prompt":"As part of my nightly eval pipeline, aggregate these LLM evaluation results from today's runs and give me back a structured analysis JSON I can store — flag anything suspicious in the warnings field. Input: {\"date\":\"2025-01-15\",\"evals\":[{\"task\":\"rag\",\"score\":0.84},{\"task\":\"summarization\",\"score\":0.79}]}."},{"title":"A/B test result consolidation","prompt":"I'm comparing two fine-tuned models and have evaluation JSON for both — please aggregate the results and tell me what the combined analysis looks like, including the evidence scope so I know how much to trust it. Data: {\"model_a\":{\"f1\":0.88},\"model_b\":{\"f1\":0.91},\"test\":\"classification\"}."}],"resultDescription":"Returns a JSON object containing a 'result' field with 'analysis_json' (structured aggregate analysis as a JSON string), 'warnings' (array of any data quality or trust cautions), 'service_id', and 'evidence_scope'. Also includes a 'status' field ('succeeded'/'failed'), full payment receipt with Base USDC transaction details, and request/execution IDs for traceability.","failureModes":["Malformed or non-JSON request_json input causes validation failure","request_json exceeds 65536 character limit returns error","Empty or minimal JSON (under 2 characters) rejected by minLength constraint","Payment failure on Base network prevents execution","Caller-supplied data flagged with trust warnings in result — data not independently verified by the service"],"whenToPreferThis":"Choose this endpoint when you need a pay-per-call, agent-native service to aggregate and structure LLM evaluation results without standing up your own eval infrastructure. Best suited for pipelines where you have raw evaluation JSON and need a structured, evidence-scoped aggregate analysis returned instantly. Prefer this over DIY aggregation when you need built-in receipt/audit trails (Base USDC, SHA256 result hash) for accountability in agentic workflows.","instructions":null,"reviewSummary":null,"reviewSummaryHighlights":null,"reviewSummaryConcerns":null,"reviewSummaryGeneratedAt":null,"activationCount":0,"lastUsedAt":null,"lastSuccessfullyRanAt":null,"lastHealthCheckAt":"2026-09-15T19:01:35.415Z","isFirstParty":false}