Build & sell

How to Build Dify Knowledge: RAG Setup Steps and Chunk Settings

・ Employee Store Operations

Summary

With Dify Knowledge, you can build AI apps that answer based on your internal manuals and FAQs. This article explains how to create a knowledge base, how to choose chunk settings and retrieval methods, and how to connect it to an app. The content was checked against the official documentation in October 2026.

Dify Knowledge is a feature that ingests documents and lets AI search them. This article explains how to create and use a knowledge base, starting with what each setting means. Screen names are given as they appear in the English interface.

Limits and default values were checked against the official Dify documentation and pricing page on October 2, 2026. They may change, so check the sources at the end before you use them.

What Dify Knowledge is

Dify Knowledge is a feature for collecting your own data and building it into AI apps. It lets the model answer based on your documents, not only on the general knowledge it learned in training. This approach is called RAG (retrieval-augmented generation).

  1. Retrieve: when a question comes in, find the most relevant parts of the knowledge base
  2. Augment: pass those parts to the model along with the original question
  3. Generate: the model writes an answer based on what it received

Documents are stored in units called knowledge bases. You can create several knowledge bases by topic or purpose and connect the right ones to each app. The difference between RAG and AI agents is explained in AI Agents vs. RAG.

There are three ways to create a knowledge base. This article covers the first, which is meant for beginners.

  • Ready-to-use: add documents and set the processing rules, and Dify handles the rest
  • Custom: build the processing steps yourself (Knowledge Pipeline)
  • External: connect to an external knowledge base through an API
Steps to create a knowledge base
  1. 1Add documentsFrom files, Notion or websites
  2. 2Chunk settingsSplit documents into small pieces
  3. 3Index and retrieval methodDecide how to search
  4. 4Wait for processingApps can use it once it finishes

Preparing documents: formats and preprocessing

On the Knowledge screen, click Create and choose a Ready-to-use knowledge base. You can add documents in three ways: upload local files, sync with Notion, or import a website. You can also create an empty knowledge base. After creation, you cannot change how documents are added.

Limits by plan (checked October 2026)
ItemSandboxProfessionalTeam
Files per upload15050
Size per fileUp to 15MBUp to 50MBUp to 50MB
Number of documents505001,000
Knowledge storage50MB5GB20GB

Images inside Word (DOCX) and Excel (XLSX) files are extracted automatically if they are JPG, PNG or GIF files under 2MB. Each extracted image is attached to its matching chunk.

Before splitting, you can also choose preprocessing to clean up the text: collapsing repeated spaces and line breaks, and removing URLs and email addresses. Removing text that has nothing to do with search improves search quality.

Take care with documents that contain personal information

Japan's Personal Information Protection Commission (PPC) issued an alert on the use of generative AI services on June 2, 2023. It asks businesses to confirm that entering prompts containing personal information stays within the stated purpose of use. Entering personal data without the person's consent, if that data is then handled for purposes other than generating the response, may violate the law. The PPC therefore asks businesses to confirm, for example, that the provider does not use the data for machine learning.

The parts found in the knowledge base are sent to the model with every question. With the High Quality method, documents are also sent to an embedding model. Before adding data such as customer lists, check how the provider of each model handles data.

How to think about chunk settings

Imported documents are split into small pieces called chunks. When a question comes in, Dify searches these chunks for relevant ones. The idea is the same as splitting a large book into chapters and paragraphs so it is easier to search.

  • Delimiter: the character that marks where to split. A blank line splits by paragraph, and a line break splits by line. The delimiter is removed when splitting, so use a character that does not appear in the text
  • Maximum length: the maximum number of characters in one chunk. Anything beyond it is split regardless of the delimiter
  • Overlap: in General mode, the number of characters shared with the neighboring chunk. At 50 characters, the last 50 characters of one chunk also start the next chunk

There are two chunking modes. You cannot change the mode after the knowledge base is created. The delimiter and maximum length can be changed later.

The 2 chunk modes

General

  • Splits all chunks with the same settings
  • Returns the matched chunks as they are
  • Suits short documents such as glossaries and FAQs

Parent-child

  • Searches with small child chunks and returns larger parent chunks
  • Finds matches precisely and passes surrounding context
  • Suits information-dense documents such as technical manuals

In Parent-child mode, you choose whether to create parent chunks per paragraph or to treat the whole document as one parent. If the whole document is the parent, only the first 10,000 tokens are processed. After you set things up, check how the text is split in Preview.

Choosing a retrieval method: vector, full-text or hybrid

Next, choose the index method: High Quality or Economical.

  • High Quality: an embedding model turns each chunk into a sequence of numbers (a vector) and searches by closeness of meaning. You cannot switch back to Economical after creation
  • Economical: searches with 10 keywords per chunk. It uses no tokens, but accuracy is lower. You can upgrade to High Quality later

With High Quality, you can choose from three retrieval methods.

Choosing a retrieval method

How do users ask questions?

You want to find content with similar meaning, even when phrased differentlyVector searchAlso works for documents in other languages
Users often search with fixed terms such as product names or model numbersFull-text searchReturns parts that contain the exact terms
Both kinds of questions come in, or you are unsureHybrid searchSearches both ways and reranks by weights or a Rerank model

Every method lets you set TopK and Score Threshold. TopK is the number of chunks returned, with a default of 3. Score Threshold is the minimum similarity to return, with a default of 0.5. For vector search and full-text search, these two settings take effect only when you use a Rerank model. Using a Rerank model costs tokens for that model.

The choice also affects storage. Vectors created with High Quality use the workspace's storage. The amount depends on the number of chunks, not the file size. Economical does not use storage.

Calling knowledge from an app

You can connect a knowledge base to any type of app. For a Chatbot, the flow is as follows.

  1. Create a Chatbot app in Studio
  2. Under Context, click Add and select the knowledge base you created
  3. Adjust the search settings in Retrieval Setting
  4. Add the Citation and Attribution feature to show sources in answers
  5. Test in the preview with questions about the documents, then publish

In a Chatflow or Workflow, use the Knowledge Retrieval node. Pass the node's results to an LLM node as context and have it write the answer. In a Chatflow, return the reply at the end with an Answer node.

Search settings work in two stages: the knowledge base and the app. The knowledge base gathers candidates, and the app reranks or narrows them. When you connect several knowledge bases, reranking by weights is available only if all of them use High Quality.

Reviewing and improving answer accuracy

When answers are off, first check whether the search is finding the right chunks.

  • Retrieval Testing: on the knowledge base screen, enter a question and test only the search. Settings changed here apply only to that test
  • Records: logs of test searches and real searches from apps are kept
  • Fix chunks: you can manually edit chunks that were split in the wrong place
  • Filter by metadata: tag documents with attributes and search only those that match a condition

The official documentation also explains how to choose a method. If users know the exact terms, set the keyword weight to 1. If the documents do not contain the same words, or questions cross languages, set the semantic weight to 1. If the documents are complex and accuracy matters, consider a Rerank model.

The basics of Dify are covered in How to Use Dify, and an example of using knowledge for support inquiries is in Using AI Agents for the Help Desk.

List your app on Employee Store

Employee Store is a marketplace where companies adopt AI agents built by developers, through a one-time purchase or a monthly plan. Any format is accepted, including Dify apps. There is no listing fee and no upfront cost. The commission is 20% of the deal amount and applies only when a deal closes. See the seller guide for details.

FAQ

Are Dify Knowledge and a knowledge base different things?
Knowledge is the name of the feature that builds your own data into AI apps. Documents are stored and managed in units called knowledge bases. One workspace can hold several knowledge bases.
How should I decide the chunk length?
There is no fixed right answer. Short chunks are found precisely but carry little context, and long chunks carry more context but lower accuracy. Test with Preview and Retrieval Testing to decide. If you need both, consider Parent-child mode.
What happens when the storage is full?
You can no longer add documents, add or edit chunks, or change chunk settings. Deleting documents you no longer need or upgrading your plan lets you add again. The documentation says that if you downgrade and exceed the limit, your data is not deleted and apps can still search it.

About the author

Employee Store OperationsThe operations team behind Employee Store, a marketplace for AI agents. We check tool features and pricing against official sources and list them at the end of each article. If you spot an error, please let us know via the contact form.

Sources

Ask AI

Ask AI if it fits your work.

Use your usual AI to explore what Employee Store offers and what to check before buying.

Opens an external AI service. Confirm pricing and deliverables on the listing page.