Instrument an AI application with DeepEval's native tracing so its behavior is visible in Confident AI.
NVIDIA RAG Blueprint — deploy, configure, troubleshoot and manage any RAG action: agentic RAG, VLM, guardrails, query rewriting…
Filesystem RAG benchmarks: corpus/, train.json, evaluate_rag.py (RAGAS quality) for a deployed NVIDIA RAG Blueprint.
Use when asked to run deep research or AI-Q research through a reachable NVIDIA AI-Q Blueprint backend.
DeepEval evaluation workflow for AI agents and LLM applications: datasets and goldens, pytest eval suites, traced evals and Confident AI…
Guides evaluation of RAG pipeline retrieval and generation quality.
Use when the user wants to search, query, extract, transcribe, quote, filter or aggregate across documents — PDFs, scanned images, Office…
Calibrate an LLM judge against human labels using data splits, TPR/TNR, and bias correction.
Design LLM-as-Judge evaluators for subjective criteria that code-based checks cannot handle.
Use to explain or reason about the foundational concepts of context engineering: the context window, attention mechanics, the U-shaped…
Help the user systematically identify and categorize failure modes in an LLM pipeline by reading traces.
Autonomous deep research from Codex via MCP.
Mem0 Platform SDK for adding persistent memory to AI applications, covering the Python and TypeScript SDKs plus framework integrations.
Use for persistent semantic memory in agent systems: cross-session retention, entity tracking, temporal validity, graph or vector…
Use when designing multi-agent systems that need context isolation, supervisor or swarm coordination, explicit handoffs, or parallel…
Integrate Mem0 into an existing repository using a goal-driven, TDD pipeline that writes failing tests before any implementation.
AI & Agents agent skills on mcprush, filed under one heading so the shelf can be read in one pass. The kinds are workflow, expertise, output format, voice & style and guardrail, and the one a skill belongs to is the fastest way to guess what it will do to your prompt.
What to look at before installing one is what it costs in context and what it takes away: a guardrail that refuses an action and a voice that rewrites a sentence are both skills, and only one of them will stop the agent doing something. Every skill here says which kind it is, what it costs in context, what it needs beside it to work and the licence it is published under, next to the publisher it came from and whether that publisher has been verified. The filters on the left narrow it further and the sort at the top decides what comes first.
A skill is not connected to anything: there is no endpoint to authorise and nothing routed through the gateway. Installing one drops a folder of instructions into the client’s own skill folder, and uninstalling it is deleting that folder. The exact line for each client is on the skill’s own page, because it differs by client rather than by skill.