Search Results for “name” – Page 32 – C4: Container, Code, Cloud & Context

Semantic Caching for LLMs: Embedding-Based Similarity and Cache Strategies

Posted on April 28, 2024

Introduction: LLM API calls are expensive and slow—semantic caching reduces both by reusing responses for similar queries. Unlike exact-match caching, semantic caching uses embeddings to find queries that are semantically similar, even if worded differently. This enables cache hits for paraphrased questions, reducing latency from seconds to milliseconds and cutting API costs significantly. This guide […]

Read more →

Function Calling Patterns: Tool Schemas, Execution Pipelines, and Agent Loops

Posted on April 18, 2024

Introduction: Function calling transforms LLMs from text generators into capable agents that can interact with external systems. By defining tools with clear schemas, models can decide when to call functions, extract parameters from natural language, and incorporate results into responses. This guide covers practical function calling patterns: defining tool schemas, handling multiple tool calls, implementing […]

Read more →

Fine-Tuning LLMs: From Data Preparation to Production Deployment

Posted on April 15, 2024

Introduction: Fine-tuning transforms a general-purpose LLM into a specialized model tailored to your domain, style, or task. While prompt engineering can get you far, fine-tuning offers consistent behavior, reduced token usage, and capabilities that prompting alone cannot achieve. This guide covers the complete fine-tuning workflow—from data preparation to deployment—using both cloud APIs (OpenAI, Together AI) […]

Read more →

Advanced RAG Patterns: Beyond Basic Retrieval

Posted on April 10, 2024

Six months ago, I thought RAG was simple: retrieve chunks, send to LLM, done. Then I built a system that needed to answer questions about 50,000 technical documents. Basic retrieval failed spectacularly. That’s when I discovered advanced RAG patterns—techniques that transform RAG from a prototype into a production system. ” alt=”Advanced RAG Patterns” style=”max-width: 100%; […]

Read more →

Testing LLM Applications: Unit Tests, Integration Tests, and Evaluation

Posted on April 8, 2024

Introduction: Testing LLM applications presents unique challenges compared to traditional software. Outputs are non-deterministic, quality is subjective, and the same input can produce different but equally valid responses. This guide covers practical testing strategies: unit testing with mocked LLM responses, integration testing with real API calls, evaluation frameworks for quality assessment, and regression testing to […]

Read more →

Function Calling Deep Dive: Building LLM-Powered Tools and Agents

Posted on April 1, 2024

Introduction: Function calling transforms LLMs from text generators into action-taking agents. Instead of just describing what to do, the model can actually do it—query databases, call APIs, execute code, and interact with external systems. OpenAI’s function calling (now called “tools”) and similar features from Anthropic and others let you define available functions, and the model […]

Read more →

Searching in

Search Results for: name

Semantic Caching for LLMs: Embedding-Based Similarity and Cache Strategies

Function Calling Patterns: Tool Schemas, Execution Pipelines, and Agent Loops

Fine-Tuning LLMs: From Data Preparation to Production Deployment

Testing LLM Applications: Unit Tests, Integration Tests, and Evaluation

Function Calling Deep Dive: Building LLM-Powered Tools and Agents