Studio: Knowledge Base September 28, 2026 19:24 Updated Table of contents: Introduction How to access the Knowledge Base Catalog Configure the catalog Configuring the Knowledge Base in the Tools tab Enhance your Knowledge Base with Preprocessing Hybrid search Best practices Blip Academy IntroductionThe Knowledge Base is a feature that centralizes documents, links, and articles so that the virtual agent can quickly access reliable information during interactions with users.It allows you to organize content into catalogs, import files, link URLs, and keep information up to date. How to access the Knowledge BaseThis is the initial screen that provides access to the Knowledge Base catalog management area. To access it, follow this path: In the contract dashboard, locate the Knowledge Base card. Click Import base. CatalogA grouping of information within the Knowledge Base. It centralizes files and URLs about the same topic, allowing the virtual agent to use this data to respond to users accurately.How to create a catalogAfter accessing the Knowledge Base, click Create catalog.This screen will only be displayed the first time you access itIn the window displayed, enter the catalog name. Click Save to finish or Cancel to exit.Add content to the catalogAfter creating a catalog, the Configure catalog screen will be displayed. There, you can add files and URLs that the AI agent will use to respond to users.Steps Click Add content. Choose between: Import file – upload a document from your computer. Add URL – link to a public web page. Import fileAllows you to upload documents directly from your computer to the catalog.Steps Select the Import file option. Accepted formats: XLSX, CSV, JSON, PDF, TXT, DOCX, PPTX, MD. Drag and drop the document or click to select it. Click Save to finish or Cancel to go back. Important points Added files serve as the basis for the agent's responses. You can combine several file types in the same catalog; however, each upload—even if it contains duplicate documents—will be added to the agent's history. Files cannot be edited: if you need to modify the content, you must delete the old file or upload a new version. The agent does not interpret content in images, so make sure all relevant information is included as text. For the agent to send media (images, audio, videos, PDFs, or links) as a response, the Knowledge Base must follow a specific model (available below in File template examples. Use the file Knowledge base example AI Agent as a reference for structuring the content). All media links must be public, ensuring that the agent can access and deliver them correctly to the user. Content type File type Recommendations Restrictions FAQ XLSX and CSV • Keeping questions and answers in the same cell or row increases Agent efficiency A column named “text” (in lowercase letters) is required for the file to be interpreted. Otherwise, an error will be displayed. FAQ DOCX and PDF • Keep questions and answers close together, on the same line or on different but consecutive lines. • Avoid images in the file or ensure that all relevant information is contained in text elements DOCX:Up to 1 million characters PDF: Up to 250,000 characters (an average of 120 to 150 pages) FAQ TXT and MD • Keep questions and answers on the same line. • If this is not possible, keep them as close together as possible—that is, on different but consecutive lines. TXT:Up to 500 KBUp to 500,000 characters Other content (manuals, reports, or guides that are not FAQs) PDF and DOCX • Remove tables of contents, if any. • Avoid images in the file or ensure that all relevant information is contained in text elements. • If the file is very long, split it into more than one file to facilitate data upload and maintenance. DOCX:Up to 1 million characters PDF:Up to 250,000 characters (an average of 120 to 150 pages) Other content (manuals, reports, or guides that are not FAQs) PPTX • Make sure the content relevant to the Agent is contained in the slides' text. Up to 30 MB Other content (manuals, reports, or guides that are not FAQs) JSON • Recommended for inserting isolated paragraphs and independent information of up to 700 characters • The “text” field must exist in this format File template examples Format Template XLSX Knowledge base example - AI Agent CSV modelo_csv_FAQ_Notificações_Ativas TXT modeo_txt_FAQ_Notificações_Ativas PDF modelo_pdf_Notificações_Ativas.pdf DOC/DOCX modelo_docs_Notificações_Ativas MD modelo_md_Notificações_Ativas PPT/PPTX modelo_pptx_Boas práticas_Templates_WhatsApp.pptx After adding a URL or file for the first time, you will be directed to the Configure catalog, screen, where you can manage all files and URLs linked to that catalog. On this screen, you can also add new items by clicking the Add content button.Note: The file import process follows the same procedure described in the Import file section.Add URLAllows you to link a public web page to the catalog, ensuring that the content is accessible to the agent without authentication.Steps Click Add content; Enter the complete address (including https://); Click Save or Cancel. Configure the catalogA screen that allows you to view and configure all items (files and URLs) in a catalog. For each piece of content, information such as name and type, date and author of the last update, and synchronization status (Active, Inactive, or Synchronizing) is displayed.Steps On the Knowledge Base screen, locate the desired catalog. Click the options menu (three-dot icon) and select Open. When you open the catalog, the complete list of files and URLs will be displayed, allowing you to view details and use the actions menu.Actions menu (three-dot icon) available for each item: Open: Displays the content of the selected item; Download: Downloads the file in its original format to your device; Update: Replaces the existing content with a new version; Restore: Reverts the content to the most recent previous version saved in Blip; Enable/Disable: Determines whether the content will be available for use by the AI agent; Delete: Permanently removes the content from the catalog. Manage linked URLsDisplays all pages associated with a public URL registered in the catalog. Only URLs accessible without authentication or blocks can be synchronized correctly.Steps For content of type URL, click Open. Enable or disable synchronization for each page. Delete pages when necessary. Use search to locate specific pages. Configuring the Knowledge Base in the Tools tabThe Tools tab allows your agent to interact with external resources and consult specific information to enrich its responses. Currently, the Knowledge Base is connected to the agent as a tool, allowing the agent to access multiple contexts in an organized way.How to add a Knowledge BaseWhen you click Add tool, a list of options will be displayed. In the Consult submenu, select the Knowledge Base option.Unlike previous versions, your agent can now have one or more tools of this type. This is useful for separating different topics (for example, one base for "Technical Questions" and another for "HR Policies"), allowing for more precise referencing in the agent's instructions.Main settingsFor each Knowledge Base tool, you must define the following fields: Name: A unique name that helps identify the tool. Description: This field is essential. Here, you explain to the agent when it should consult this base and what it will find there. Example: "Use this tool to answer questions about prices, plans, and payment methods." Selecting content (catalogs)For the tool to have information to consult, it needs Catalogs. Catalogs centralize your files and URLs. Click Add catalog. In the window that opens, you can select existing catalogs or click Create catalog to be redirected to the content management area. Practical tip: You can select an entire catalog or only specific content within it by checking the checkboxes. This gives you complete control over what each tool can "read."Optimization and query settingsTo ensure that the agent responds quickly and economically, the platform uses RAG (Retrieval-Augmented Generation) technology. This means that, instead of reading all documents at once, the agent searches only for the excerpts most relevant to the user's context.In the Query settings section, you can adjust: Returned excerpts (Chunks): Defines how many information "pieces" the agent receives per query. Recommendation: The default value is 3. Adding more excerpts may help the agent respond better, but it consumes more tokens and may exceed the model's context window. Enhance your Knowledge Base with PreprocessingPreprocessing optimizes your Knowledge Base documents before they are made available to Studio agents. This process improves information quality, search accuracy, and the efficiency of generated responses.We offer three types of preprocessing that can be enabled according to your needs. 1. Optimization (Noise Removal)What is it?Optimization is an automatic cleaning process that removes unnecessary "noise" and formatting from your document text. The goal is to standardize the content, ensuring that AI focuses only on relevant information.How does it work?This processor applies a set of rules to refine the text. Based on its default configuration, it performs the following actions on each section of the document: Unicode correction: Repairs characters that were corrupted or incorrectly encoded. HTML removal: Eliminates all HTML tags (for example, <div>, <p>, <span>) that may be present in documents extracted from the web. Space normalization: Removes multiple spaces, tabs, and excessive line breaks, replacing them with a single space or line break, respectively. Practical example:Original text:<p>The meeting will be on Tuesday.Check the main topic.</p>Optimized text:The meeting will be on Tuesday.Check the main topic. 2. Indexing (AI Summaries and Tags)What is it?Indexing uses an AI Agent to enrich each fragment (chunk) of your document with a concise summary and relevant keywords. This creates a "semantic index" that dramatically improves the system's ability to find the exact information the user is looking for.How does it work?An AI agent specialized in knowledge synthesis analyzes each section of text and generates: Summary: A short summary (2–3 sentences) explaining the specific topic of that section. Keywords: A list of 5 to 8 essential terms, entities, or technical jargon found in the text. The process is performed in the same language as the original document to maintain consistency. Practical example: Chunk text:"Article 14 of the service contract stipulates that the contracted party must notify the contracting party 30 days in advance of any scheduled interruption. Failure to comply with this clause will result in financial penalties, as detailed in Appendix B." Indexing result: Summary: "This section details Article 14's clause requiring 30 days' advance notice for service interruptions. Failure to provide notice results in financial penalties." Keywords: "service contract, Article 14, advance notice, scheduled interruption, clause, financial penalties, Appendix B" When should it be used?For documents that function as a collection of independent information, where each section has value on its own. Ideal for: Question-and-answer bases (FAQs): Where each question-and-answer pair is a "fact" that must be found independently. Tabular or divided content: Documents in which information is already segmented into blocks, such as a spreadsheet with product descriptions or a list of internal policies. Articles or blog posts: Where each paragraph or section addresses a specific subtopic that can be summarized to facilitate search. Main benefitCreates a rich "index" that allows the search engine to find specific sections with high precision, even when the user's search uses synonyms or related terms.CostConsumes AI tokens according to the size of the document. 3. Contextualization (Global Semantic Reading with AI)What is it?Contextualization is the most advanced preprocessing method. It uses an AI Agent to perform a "global semantic reading" of the entire document. For each text fragment, AI describes where it fits within the document's overall context, acting as a "GPS" for the search engine.How does it work?Unlike Indexing, which focuses on the content of the section, Contextualization focuses on its location and purpose. The AI agent reads the entire document to understand its structure (chapters, sections, and flow of ideas), then generates a short sentence (15–25 words) for each section describing its contextual role.Practical example: Document: "Company IT Security Manual" Chunk text:"All employees must use passwords with at least 12 characters, including uppercase letters, lowercase letters, numbers, and symbols." Contextualization result: Generated context: "'Password Policy' section of the manual, specifying complexity requirements for employees' access credentials." When should it be used?For long, continuous, and structured documents in which the position of information within the whole is crucial to its meaning. Ideal for: Technical manuals and user guides: Where it is important to know whether a section is in the "Installation," "Troubleshooting," or "Advanced Settings" section. Contracts and legal documents: Where the context of a clause (for example, "Termination Clause" or "Penalties Appendix") is essential. Scientific articles and research reports: Where the structure (Introduction, Methodology, Results, Conclusion) gives meaning to each part of the text. Main benefitActs as a "GPS" for search, telling the system not only what is in the section, but where it fits in the document's flow of information.CostConsumes AI tokens according to the size of the document. Comparison table Option Ideal for... Main benefit AI cost Optimization All file types, especially "messy" ones. Ensures the quality and consistency of the base text. No Indexing Segmented content (FAQs, tables, short articles). Makes each section findable by its specific content. Yes Contextualization Long and structured documents (manuals, contracts). Locates information within the document's overall structure. Yes Golden rule If your documents are a collection of facts, where each can be read independently, choose Indexing. If your documents tell a story or follow a logical structure, where context is key, choose Contextualization. Optimization is like tidying the house before decorating it: use it whenever possible. How to use preprocessing in StudioNow, when you click “Import files” inside the catalog, a side menu will open:In this side menu, you can upload multiple files and choose which preprocessing to apply to each one. Simply select the checkbox for the desired preprocessing and then click save:The bases will be loaded and preprocessed. Processing time may vary according to the size of the base and the selected settings. Hybrid searchHybrid search is a feature that optimizes the user's journey to quickly resolve questions by combining natural language understanding with lexical searches for precise terms. This feature automatically adjusts weights. The system establishes the ideal balance between conversational context and searches for exact terms in active bases. The efficiency of this search also depends heavily on the number of chunks configured in the interface. The number of chunks returned by your Knowledge Base tool is just as important as the quality of the stored data.Important information The interface has a default setting of 3 returned chunks. You must have an active and structured Knowledge Base. Preferably, chunks should contain between 300 and 800 characters. The recommended size is 1 to 3 paragraphs. The base should be free of redundancies. The base should be well categorized. The base should be updated regularly. How it worksStandard search systems can understand the general meaning of a user's question using natural language. The system may lose precision when the user uses highly specific terms. The main losses in precision occur with numbers, codes, and exact names of functions present in active bases. To mitigate these limitations, hybrid search combines two search approaches simultaneously: Understanding the question's context: Focuses on the semantic understanding of the user's intent. Searching for exact terms: Focuses on lexical search to retrieve keyword accuracy. Automatic actions of Dynamic Hybrid SearchThe platform provides the Dynamic Hybrid Search version. This version's main technical differentiator is the automatic adjustment of the relevance weight between context and exact terms. The balance depends strictly on the type of question submitted: More conversational question: The system prioritizes context understanding through the semantic search already available on the platform. More technical or specific question: The system prioritizes exact terms through the newly added lexical search. Advantages: Search improves the retrieval of technical answers. Search places special emphasis on finding exact numbers and names. The transition prevents answers to natural-language questions from getting worse. The update increases the consistency and robustness of the entire search experience within an active Knowledge Base. How to optimize the number of returned chunksThe hybrid search feature works entirely automatically in its balancing. The only technical setting required on the platform is the number of configured chunks. By default, the number of chunks configured in the interface is 3. Depending on the size of the base, this value may be too small and prevent the intelligence from working properly. The number of configured chunks represents the number of pages the agent can consult before formulating a response.If the base is very large (approximately 150 chunks) and the setting is too low (such as 5 chunks), the agent will experience the following failures: Incomplete retrieval: Semantic search finds the relevant documents, but only 3 or 4 actually reach the agent. Contextual hallucination: Without enough context, the agent begins to invent answers or make incorrect assumptions. Loss of nuance: Complementary information contained in unretrieved chunks is lost by the system. Inconsistency: The same question may generate completely different answers depending on the chunks returned in that request. Recommendation matrix based on base sizeThe number of chunks must keep pace with the size of the base. The larger the base, the more chunks must be returned. Follow the guidelines below to align the interface settings: Suitable base (50 to 150 records): Configure between 5 and 15 chunks for standard flows. Suitable complex base: Configure between 8 and 20 chunks for complex flows, achieving maximum coverage at a low cost. Medium base (150 to 250 records): Configure between 15 and 25 chunks for standard flows. Complex medium base: Configure between 20 and 35 chunks for complex flows, ensuring a balance between response accuracy and latency. Large base (250 to 500 records): Configure between 25 and 50 chunks for standard flows. Complex large base: Configure between 35 and 65 chunks for complex flows, since less relevant documents begin to appear in the search. Very large base (more than 500 records): Configure up to the technical maximum limit of 500 chunks. Base division: For scenarios that exceed this capacity, it is recommended to divide the knowledge into sub-bases categorized by Support, Finance, and Products. Multiple agents: As an alternative, use the Model Router with multiple agents working simultaneously. Configuration validation testsAfter applying the calculations, implement these three security assessments: Coverage test: Gather 10 questions representative of your flow. Coverage assessment: Run each question and confirm that the returned chunks cover more than 80% of the ideal answer. Hallucination test: Run exactly the same question 5 consecutive times. Inconsistency assessment: If the answer varies significantly, there is high inconsistency, indicating insufficient chunks. Inconsistency correction: Increase the setting by 5 to 10 chunks and repeat the assessment. Relevance score test: Access the Search tab. Relevance assessment: Analyze the score identified as relevance_score of the last returned chunk. Score lower than 0.50: Increase the number of chunks. Score between 0.50 and 0.65: Consider increasing the number of chunks. Score higher than 0.65: The setting is adequate and should be maintained. Pay attention to technical cost and latency limitations Using a high number of chunks significantly increases the volume of data consumed. Using 15 chunks consumes an average of 450 input tokens. Using 15 chunks adds 150 to 200 ms of latency. The maximum volume of 500 chunks consumes approximately 15,000 input tokens. The maximum volume of 500 chunks generates high latency ranging from 2,000 to 2,500 ms. Evaluate the maximum response time allowed in your flow's latency SLA. If the maximum tolerated time is 500 ms, limit the setting to a maximum of 50 chunks. If the maximum tolerated time is 1 second, limit the setting to a maximum of 100 chunks. Best practices Keep content up to date. Use clear, standardized names to make searches easier. Review URLs periodically to avoid broken links. Blip AcademyWant to learn how Blip Studio works and how to work with it? Access Blip Academy and learn for free. For more information, access the discussion about this topic in our community or the videos on our channel. 😃 Related articles Studio: Getting Started - Basic Settings How to configure and send WhatsApp Active Messages in Blip Desk How to Schedule a Message with the Scheduler Extension How to Use Variables in Blip Desk Canned Responses How to import/export a knowledge base