PepoChat
Knowledge baseNotionGoogle Drive

How to Build an AI Support Chatbot From Notion and Google Drive Docs

Connect Notion pages and Google Docs, Sheets and Drive files as live knowledge sources for a support chatbot, structure them for retrieval, and sync daily.

PepoChat TeamPublished Last verified 16 min read
Stacks of paper folders with yellow tabs piled high on an office desk under fluorescent lights

Short answer

You build an AI support chatbot from Notion and Google Drive by connecting them as knowledge sources rather than exporting PDFs. In PepoChat, each connector syncs the pages or files you choose, converts them to text the agent can search, re-fetches only what changed once a day, and counts as a single knowledge source with up to 500 pages or files. Nothing is fine-tuned: the agent retrieves the relevant passage at answer time, so an edit in Notion or Docs reaches customers within a day.

If your product documentation, policies and internal how-tos live in Notion or Google Drive, you have probably tried the export route at least once: download a PDF, upload it to a chatbot, and discover three weeks later that the bot is confidently quoting a refund policy you changed. The export is a snapshot; the source keeps moving.

This guide is for founders, support leads and operations people who keep the company's knowledge in Notion pages, Notion databases, Google Docs or Google Sheets, and who want a website chatbot that answers from that content without a second copy to maintain. It explains what "training" a chatbot on Notion actually means, how the two connectors in PepoChat work step by step, how to structure your pages so retrieval finds the right fact, and what to keep out of the shared set.

The product specifics are PepoChat's. The structuring and hygiene advice applies to any retrieval-based support agent that reads from a shared workspace.

What does "training" a chatbot on Notion or Google Drive mean?

It does not mean the model learns your documents. Retrieval-augmented generation (RAG) is the technique of searching your own content for the passages relevant to a question and handing those passages to the language model at the moment it writes the reply. Wikipedia's retrieval-augmented generation article describes it as a technique that enables large language models to retrieve and incorporate new information from external data sources, and notes that it reduces the need to retrain the model with new data.

That distinction matters for a Notion or Drive workflow in two practical ways. First, updates are cheap: because nothing is retrained, a changed page only needs to be re-fetched and re-indexed, not re-learned. Second, the quality of the answer depends almost entirely on whether the right passage exists and can be found, which is why the structuring section below matters more than any setting in the dashboard. The longer explanation, with an analogy and a glossary, is in RAG for customer support, explained for non-engineers.

In PepoChat, every reply starts with a search over the workspace's private knowledge base, and when nothing relevant comes back the agent hands off to a person rather than guessing. A knowledge source is one thing you added to that knowledge base: a crawled website, an uploaded file, or a connected app. A Notion connection or a Google Drive connection counts as one source however many pages it holds.

Four ways to get documents into a support chatbot

Connectors are not the only way in, and for some content they are not the best way. Here is how the four options in PepoChat compare.

MethodWhat it takes inKeeps itself currentCounts asCeiling
Notion connectorPages and databases you share with PepoChatYes, changed pages re-fetched once a day, plus Sync now1 source500 pages per connection
Google Drive connectorDocs, Sheets, Slides and PDF, DOCX, TXT, Markdown, CSV, JSON files you pickYes, changed files re-fetched once a day, plus Sync now1 source500 files per connection; files over 5 MB skipped
File uploadPDF, DOCX, Markdown, HTML, TXT, CSV, JSON, imagesNo, re-upload a file with the same name to replace it1 source per file25 MB per upload
Website crawl or sitemapPublic pages on your siteYes, re-fetched weekly, only changed pages re-indexed1 source per crawl50 pages and 3 levels deep per crawl

Text from every method counts toward the plan's knowledge-base text cap: 20 MB on Free, 200 MB on Starter and 1 GB on Growth. The Free plan allows five knowledge sources, and because a connection is one source, a whole Notion wiki fits in one slot. Details of the plan limits are on the pricing page and in the manage sources documentation.

How to connect Notion to an AI support chatbot

The Notion connector lives on the Connected apps tab of the Knowledge Base page. The whole setup takes a few minutes; the first sync takes as long as the number of pages you share.

Stacks of paper folders with yellow tabs piled high on an office desk under fluorescent lights
Exported PDFs are a snapshot of your Notion workspace. The connector reads the live pages instead.
  1. Open Knowledge Base, click Add New and choose Connected apps

    Connected apps is the last tab of the Add New dialog. Click Connect Notion. If the card reads "Not available yet", the connector has not been switched on for your deployment; availability may be rolling out, and everything else on the tab works as described once the card shows the Connect button.

  2. Share pages on Notion's screen

    Notion asks which pages to share with PepoChat. Only the pages and databases you share here are ever visible to the connector. Notion's own help page on connections describes the same model: when you install a connection through a partner platform, you select which pages it can access during authentication, and you can add more later from the page menu or from Settings under Connections. Click Allow. You land back on the Knowledge Base with a confirmation toast, and a Notion row appears above the table with the badge Connected, "0 pages" and "Never synced".

  3. Choose what to sync

    Open Add New and the Connected apps tab again. The Notion card shows your account and "Everything shared with PepoChat". Click Choose pages to get a searchable checklist of the shared pages and databases, with databases tagged. Ticking a page includes everything nested under it. Leaving every box empty means everything you shared will be synced. Click Save and sync.

  4. Watch the first sync

    The imports panel at the top of the page shows a Connected app card moving through "Finding pages", "Importing" with a running count, and "Completed". The connection row updates to the page count and "Last synced less than a minute ago". You can Cancel a running sync from the panel; the connection returns to Connected.

What the agent sees from Notion

  • Pages are converted to Markdown. Headings, lists, tables, toggles, callouts and code blocks are kept, so the structure you wrote survives.
  • Databases become one page per row: the row's title, a table of its properties, then the row's page body. A "Known issues" database with 40 rows becomes 40 short, searchable pages.
  • Each synced page appears in the Knowledge Base table with a NOTION badge, its title, its size and a link back to the page in Notion, filed under the category "Notion".
  • Images and files embedded in a page are not read. The surrounding text is. If a fact lives only in a screenshot, write it out as a sentence next to the image.

Keeping Notion in sync

Once a day, PepoChat checks the "last edited" time of every page in the selection and fetches only the ones that changed. Unchanged text is never re-indexed. Pages you delete, unshare or untick disappear from the knowledge base at the next sync. Sync now, on the connection row or the card, runs the same check immediately, and after a single edit the imports panel shows exactly that one page being imported.

Two limits to know. A connection holds up to 500 pages; past that, the pages already synced are kept and nothing new is added until you narrow the selection, and the connection row tells you so. And if you share more than 2,000 items with PepoChat, each sync checks only the 2,000 most recently edited; pages outside that window are kept as they are, neither refreshed nor removed. Share fewer pages to keep the sync exact.

One import runs at a time per workspace, and website crawls share the same queue. If you save a selection while a crawl is running, the sync starts automatically when the crawl finishes.

How to connect Google Drive to an AI support chatbot

The Google Drive connector follows the same shape but uses Google's file picker, which changes what "sharing" means: you choose individual files, and PepoChat never sees anything you did not pick.

  1. Open Knowledge Base, click Add New and choose Connected apps

    Click Connect Google Drive. The same "Not available yet" note applies if the card shows it.

  2. Allow access on Google's screen

    Google asks for permission to access only the specific Drive files you use with the app. There is no request to read your whole Drive. Click Allow. Back on the Knowledge Base, a Google Drive row appears above the table showing your Google email; its Sync now button stays disabled until you pick files. If Google's consent screen comes back without the Drive permission ticked, PepoChat tells you to tick it and try again.

  3. Pick files

    Open Add New and the Connected apps tab, then click Choose files on the Google Drive card. Google's picker opens. Select as many files as you like; folders are shown but cannot be picked, so open a folder and select the files inside it. Click Select. A toast reports how many files were added and how many of the 500-file allowance are used, and the imports panel shows the sync.

  4. Add more later

    Click Choose files again whenever you like. New picks are added to the earlier ones and only the new files are fetched.

How each Google file type is read

FileHow it is readWatch out for
Google DocsExported as Markdown with headings keptNothing special; this is the best-behaved type
Google SheetsExported as CSVOnly the first sheet of a spreadsheet is synced. Put the facts the agent needs on the first tab
Google SlidesExported as plain textSpeaker notes and layout are lost; a Doc is usually better for support content
PDF, DOCX, TXT, Markdown, CSV, JSON uploads in DriveDownloaded and processed like an uploaded document; scanned PDF pages are read by a vision modelFiles over 5 MB are skipped
Google Drawings, Forms, Sites, shortcuts, imagesSkippedThe imports card counts them as failed; the other files sync normally

Each synced file becomes a row with a GDRIVE badge, the file's name, its size and a link that opens the file in Drive, under the category "Google Drive".

Keeping Drive in sync

Once a day, only files whose modification time changed since the previous sync are fetched again. A file you move to the trash in Drive, or that is no longer shared with you, disappears from the knowledge base at the next sync. Sync now runs the same check immediately.

There is no unpick button. To remove one file for good, trash or unshare it in Drive. To start over, Disconnect and connect again with fewer files. Deleting a file's row from the Knowledge Base table only removes it until the next sync brings it back. A pick that would push a connection past 500 files is refused with a message telling you the count it would reach.

Because removal is driven by Drive's own sharing state, it helps to know how Google handles access. Google's sharing help page sets out the viewer, commenter and editor roles and notes that folder permissions apply to the files inside them. The connector needs only to be able to read a file, and the Google account you connect must still have access to it at each sync.

How to structure Notion pages and Google Docs so the chatbot answers well

Retrieval finds passages, not documents. The agent pulls the few chunks of text that best match the visitor's question and answers from those. Everything in this section follows from that single fact.

A man in a grey turtleneck presents a plan on a flipchart to three colleagues seated around a wooden table
Write pages the way a colleague would explain them: one topic, plain statements, the question in the heading.

One topic per page. A single 6,000-word "Everything about billing" page is retrieved in pieces, and the piece that matches "do you charge VAT" may not contain the sentence that answers it. Five pages of a few hundred words each, one per billing topic, are retrieved whole and answer cleanly. Notion databases are a natural fit here because each row becomes its own page.

Put the question in the heading. Headings survive the conversion to Markdown, and a heading that reads "Can I change my plan mid-month?" matches the way visitors ask far better than "Plan changes". The same rule that PepoChat's best practices guide gives for Q&A pairs applies to Notion headings.

Write statements, not marketing. "Orders ship within 2 business days. Express delivery costs $9." is quotable. "Lightning-fast shipping" is not. The agent paraphrases what it reads, so the text has to contain the fact.

Keep the facts on the first tab of a Sheet. Only the first sheet syncs. A pricing matrix on tab three is invisible. Either move it to tab one or split the spreadsheet into one file per table.

Write out what images show. Embedded images are skipped in both connectors. A screenshot of the settings page with an arrow on the right button tells the agent nothing; a sentence saying "The export button is under Settings, then Data" tells it everything.

Keep one version of the truth. If an old pricing page and a new one are both shared, the agent may retrieve either. Untick the old page or archive it in Notion. The same goes for a Doc and its earlier draft in Drive: pick one.

Use toggles and callouts freely. Both are kept in the conversion, so a Notion page built from toggles is fine. What matters is that the answer text is inside the page, not linked from it: the connector reads the page you shared, not the pages it links to unless you shared those too.

Use Q&A pairs for the last mile. When the agent gets one question wrong, or cannot answer it, you do not need to restructure a page. Add a Q&A pair with the question the way the visitor asked it and the answer you want. Pairs do not count as knowledge sources, so they never touch the Free plan's five-source limit.

What to keep out of the shared pages

Anything you share is answerable. That is the whole point of the connector and also its main risk, because Notion workspaces mix customer-facing documentation with pages that were never meant to leave the building.

  • Internal candour about customers or accounts. "Client X is difficult, escalate slowly" belongs in a CRM note, not a page the support agent can read back to a verified visitor.
  • HR, hiring and compensation pages. Salary bands, interview scorecards and performance notes are the classic accidental share when someone ticks a parent page and everything nested under it comes along.
  • Unfinished drafts and proposals. A "Pricing 2027 proposal" page reads to the agent exactly like the live pricing page. Keep drafts in a section you do not share, or move them out of the shared tree until they are approved.
  • Credentials and keys. Some teams keep API keys and vendor logins in a Notion database. Never share that database. The agent would not print a key on purpose, but the text would still be stored in your knowledge base.
  • Meeting notes and retrospectives. They are full of half-decisions and names. If you want the agent to know what was decided, write the decision on the relevant documentation page instead.

The practical habit that prevents all of this: keep a top-level "Customer docs" page in Notion, share only that page with PepoChat, and let nesting do the rest. In Drive, keep the files you sync in one folder so the picker is a short list rather than a hunt through the whole account. A quarterly look at the Knowledge Base table, which lists every synced page by title, catches anything that slipped in.

How to test the chatbot before customers see it

Connecting is not the same as verifying. Three checks, in order, catch most problems.

  1. Ask the ten questions you get most. Open the Test the agent drawer from the Knowledge Base page and ask real questions in the words visitors use. Test chats are free and stay out of analytics. If an answer is wrong, click through to the synced page listed in the table and read what the agent read; the fix is nearly always on the page, not in the settings.
  2. Ask something the pages do not cover. The right behaviour is an honest "I don't know" followed by an offer to hand the conversation to a person. If the agent invents an answer instead, a page in the shared set is misleading it; the guide to stopping chatbot hallucinations walks through how to find which one.
  3. Edit a page and wait for the sync. Change one sentence in Notion or a Doc, click Sync now, watch the imports panel show that single page importing, then ask the question again. This proves the loop works end to end before you rely on it.
An open notebook with handwritten notes and green sticky notes lies beside a laptop keyboard on a sunlit desk
Test with the questions customers actually ask, then fix the page the agent read rather than the agent.

Troubleshooting Notion and Google Drive sync

The connection row above the Knowledge Base table shows one of four statuses. Most problems are visible there.

SymptomWhat it meansWhat to do
Badge reads Sync failedThe last run failed; the row shows a short reasonFix the cause it names, then click Sync now
Badge reads Access revokedNotion or Google no longer accepts the connection, for example after you removed PepoChat under Notion's Connections or your Google account's third-party accessClick Reconnect; your page selection or file picks are kept
"Your plan's knowledge-source limit is reached" when connectingThe Free plan's five sources are used and a connection needs one slotDelete a source you no longer need, or upgrade under Plans and Billing, then connect again
Connection row says the 500-page or 500-file cap is reachedPages already synced are kept; nothing new is addedNarrow the selection in Choose pages, or for Drive, disconnect and reconnect with fewer files
A page is synced but the agent does not answer from itThe passage that answers the question is not on the page, is inside an image, or is on a later Sheet tabWrite the fact as a sentence, move it to the first tab, or add a Q&A pair
A page you deleted from the table came backDeleting a row only removes it until the next syncUntick it in Choose pages or unshare it in Notion; trash or unshare it in Drive
Text cap warning on the Knowledge Base pageSynced text counts toward the plan's 20 MB, 200 MB or 1 GB cap like uploadsUntick large pages you do not need, or upgrade
"Google didn't grant Drive access"The Drive permission was not ticked on Google's consent screenReconnect and tick the Google Drive permission

If none of these match, the troubleshooting page covers wrong and missing answers in general.

Privacy: where synced content goes and how to remove it

Content synced from Notion or Google Drive is stored in your workspace's knowledge base exactly like an uploaded document: it is converted to text, indexed for retrieval, and isolated to your organisation. When the agent answers, the passages it retrieved are sent to the model provider along with the conversation; the rest of the page is not. PepoChat's privacy and GDPR notes describe what to put in your own privacy policy, including naming the connected tools whose data the agent reads.

Removal is immediate and complete at the connection level. Open Add New, Connected apps, click Disconnect on the card and confirm. The connection row and every page or file it synced are removed from the knowledge base at once. For Google Drive, PepoChat also asks Google to revoke its access to the account. To switch to a different Notion workspace or Google account, disconnect first and connect again.

At the page level, removal follows the source: untick or unshare in Notion, trash or unshare in Drive, and the page is gone at the next sync or sooner with Sync now. That is the right mental model for the whole feature. Your Notion workspace and your Drive folder are the source of truth; the knowledge base is a searchable mirror of exactly the part you chose to share, and it stays a mirror rather than becoming a second copy to maintain.

What to do next

If your documentation is already in Notion or Drive, connect it before you upload anything else; a single connection fills one of the five free knowledge sources and covers hundreds of pages. Then run the three tests above with real customer questions. For content that lives on your public website rather than in a workspace, the guide to training a chatbot on your website and PDFs covers crawls and sitemaps, and the installation guide gets the widget onto your site with one script tag. Everything here is on the free plan when you create a workspace, with no card needed to start.

Frequently asked questions

Can you train an AI chatbot on Notion pages?
Yes, but training here means retrieval, not fine-tuning. A connector syncs the Notion pages and databases you share, converts them to searchable text, and the agent pulls the relevant passage at answer time. In PepoChat a Notion connection holds up to 500 pages, re-fetches only changed pages once a day, and counts as a single knowledge source.
How does a chatbot read Google Docs, Sheets and Slides?
Google Docs are exported as Markdown with headings kept, Sheets as CSV from the first tab only, and Slides as plain text. PDF, DOCX, TXT, Markdown, CSV and JSON files stored in Drive are processed like uploads, with scanned PDF pages read by a vision model. Drawings, Forms, Sites, shortcuts and images are skipped, as are files over 5 MB.
Does the chatbot see my whole Notion workspace or Google Drive?
No. Notion asks which pages to share with the connector during authorisation, and only those pages and anything nested under them are visible. Google Drive uses Google's file picker, so only the files you select are accessible and there is no request to read the whole Drive. Sharing a parent page in Notion does include everything nested beneath it, so check before ticking.
How often does the chatbot update when I edit a page?
Both connectors check once a day and fetch only pages or files whose last-edited time changed, so unchanged text is never re-indexed. Clicking Sync now runs the same check immediately, and after a single edit the imports panel shows that one page importing. Pages you untick, unshare, trash or delete disappear from the knowledge base at the next sync.
Why does the chatbot not answer from a page I synced?
Usually because the answering sentence is not actually on the page: the fact is inside an image, which connectors skip, on a later Sheet tab, which is not synced, or buried in a long page where the matching passage lacks the answer. Write the fact as a plain sentence under a question-shaped heading, or add a Q&A pair, which does not count as a knowledge source.
How do Notion and Google Drive connectors count toward PepoChat's plan limits?
A connection is one knowledge source however many pages or files it holds, so a whole wiki fits in one of the Free plan's five sources. The synced text counts toward the plan's knowledge-base text cap of 20 MB on Free, 200 MB on Starter and 1 GB on Growth. Disconnecting removes the connection and everything it synced immediately, and for Drive also revokes Google access.

Try this on your own site in ten minutes

PepoChat includes every feature on the free plan — 200 AI replies and 5 knowledge sources a month, no credit card.