---
title: "Creating Knowledge Graphs from Unstructured Data - Developer Guides"
url: https://memory.wiki/HAPYFIou
updated: 2026-10-09T20:35:32.692Z
hub: https://memory.wiki/hub/pratofeito
bundle_count: 1
concept_count: 12
source: "url:neo4j.com"
---
# Creating Knowledge Graphs from Unstructured Data - Developer Guides

# Creating Knowledge Graphs from Unstructured Data

Creating graph structures from unstructured text using [NLP techniques like Spacy, Stanford NLP, OpenNLTK](https://neo4j.com/blog/machine-learning/accelerating-towards-natural-language-search-graphs/) etc. has been a long-standing possibility. Those usually required fine-tuning of the NLP models on the specific domain and language of the text to be processed and substantial NLP skills to get the best results.

The rise of prompt-able modern day Large Language Models (LLMs) with strong language skills across many languages and domains, extended the capabilities massively. The LLMs can be instructed with detailed prompts (instructions, examples, schema, existing entities, output formatting) to extract, and de-duplicate entities and relationships from unstructured text, images and audio fragments. The extracted information is returned in a structured JSON format and can be stored in a graph database like Neo4j and linked back to the source documents and additional metadata.

This allows to build Knowledge Graphs (KGs) from unstructured data and integrate them with existing structured data in Neo4j. Those Knowledge Graphs can then be used for a variety of applications, including [GraphRAG](https://graphrag.com) (Retrieval Augmented Generation) applications, where the Knowledge Graph is used to retrieve relevant contextual information to augment the LLM’s responses.

![0*qgT2hBiA3DA1Y3qu](https://miro.medium.com/v2/resize:fit:1400/format:webp/0*qgT2hBiA3DA1Y3qu.png)

[![llm vectors unstructured](https://cdn.graphacademy.neo4j.com/assets/img/courses/banners/llm-vectors-unstructured.png)](https://graphacademy.neo4j.com/courses/llm-vectors-unstructured/?category=generative-ai)

## [](#_neo4j_product_capabilities)Neo4j Product Capabilities

### [](#_graphrag_python_package)GraphRAG Python Package

The Neo4j Python GraphRAG package offers a comprehensive Pipeline for unstructured document processing, graph schema based entity extraction, resolution and storage.

The `SimpleKGBuilder` provides an easy way to get started, for full configurability, please use

```python
from neo4j_graphrag.experimental.pipeline.kg_builder import SimpleKGPipeline

ENTITIES = [
    "Person",
    {"label": "Company", "description": "Company or organization"},
    {"label": "Location", "properties": [{"name": "city", "type": "STRING"}]},
]
RELATIONS = [
    "LOCATED_AT",
    {
        "label": "COMPETES_WITH",
        "description": "Used for competitor relationships between companies",
    },
    {"label": "WORKS_AT", "properties": [{"name": "fromYear", "type": "INTEGER"}]},
]

kg_builder = SimpleKGPipeline(
    llm=llm, # an LLMInterface for Entity and Relation extraction
    driver=neo4j_driver,  # a neo4j driver to write results to graph
    embedder=embedder,  # an Embedder for chunks
    from_pdf=True,   # set to False if parsing an already extracted text
    entities=ENTITIES,
    relations=RELATIONS,
    potential_schema=POTENTIAL_SCHEMA, # a optional list of node-relationship-node schema entries
)
await kg_builder.run_async(file_path=str(file_path))
```

![kg builder pipeline](https://neo4j.com/docs/neo4j-graphrag-python/current/_images/kg_builder_pipeline.png)

A full Knowledge Graph (KG) construction pipeline requires a few more components:

-   **Data loader:** extract text from files (PDFs, …).
    
-   **Text splitter:** split the text into smaller pieces of text (chunks), manageable by the LLM \* context window (token limit).
    
-   **Chunk embedder (optional):** compute the chunk embeddings.
    
-   **Schema builder:** provide a schema to ground the LLM extracted entities and relations and obtain \* an easily navigable KG.
    
-   **Lexical graph builder:** build the lexical graph (Document, Chunk and their relationships) \* (optional).
    
-   **Entity and relation extractor:** extract relevant entities and relations from the text.
    
-   **Knowledge Graph writer:** save the identified entities and relations.
    
-   **Entity resolver:** merge similar entities into a single node.
    

Learn more in the documentation:

-   [Neo4j GraphRAG Python Package Overview](https://neo4j.com/developer/genai-ecosystem/graphrag-python/)
    
-   [User Guide: Knowledge Graph Builder](https://neo4j.com/docs/neo4j-graphrag-python/current/user_guide_kg_builder.html)
    
-   [User Guide: Pipeline](https://neo4j.com/docs/neo4j-graphrag-python/current/user_guide_pipeline.html)
    
-   [GraphAcademy Course](https://graphacademy.neo4j.com/courses/llm-vectors-unstructured/?category=generative-ai)
    

Those generated Knowledge Graphs can then be used with the [same Python package, for building GraphRAG powered](https://neo4j.com/docs/neo4j-graphrag-python/current/user_guide_rag.html) GenAI applications.

## [](#_prototyping_and_demonstrations)Prototyping and Demonstrations

### [](#_llm_knowledge_graph_builder)LLM Knowledge Graph Builder

The [LLM Knowledge Graph Builder](https://neo4j.com/labs/genai-ecosystem/llm-graph-builder/) is an easy to use tool for building Knowledge Graphs from unstructured text.

It uses the LangChain LLMGraphTransformer under the hood, but provides a simple UI to extract entities and relations from uploaded PDFs, Office Documents, Web pages or YouTube video transcripts. It extracts first a lexical graph (Document, Chunk and their relationships) and then uses the chosen LLM to extract entities and relations. Optionally it enriches the graph community summaries and embeddings. You can specify a graph schema to guide the extraction process and visualize the extracted graphs.

![build kg genai e1718732751482](https://dist.neo4j.com/wp-content/uploads/20240618104511/build-kg-genai-e1718732751482.png)

To interact with the extracted information, you can compare a number of different retrievers including GraphRAG, Vector, Hybrid and Text2Cypher. For each answer you can inspect the source of the information and the retrieved context that was used to generate the answer. It also allows to run an RAGAs evaluation across the retrieval results.

Some of the internal details are documented in [LLM Graph Builder - Knowledge Graph Extraction Challenges](https://neo4j.com/blog/developer/knowledge-graph-extraction-challenges/).

### [](#_ai_assistant_with_neo4j_mcp_server)AI Assistant with Neo4j MCP Server

A quick way of experimenting interactively with Knowledge Graph extraction is to use a AI assistant like Claude, ChatGPT or an MCP enabled IDE like VS Code, Cursor, Windsurf to process text or discussions interactively and prompt the LLM to extract Entities and Relationships.

Previously those extracted graph entities would have to be imported through an intermediate format like JSON or CSV (e.g. via the Data Importer) or via Cypher.

Now they can be automatically saved to a connected graph using the [mcp-neo4j-cypher](https://github.com/neo4j-contrib/mcp-neo4j/tree/main/servers/mcp-neo4j-cypher) or [mcp-neo4j-memory](https://github.com/neo4j-contrib/mcp-neo4j/tree/main/servers/mcp-neo4j-memory) MCP servers.

![employee create entities and relations](https://github.com/neo4j-contrib/mcp-neo4j/raw/main/servers/mcp-neo4j-memory/docs/images/employee_create_entities_and_relations.png)

## [](#_ecosystem_integrations)Ecosystem Integrations

Neo4j integrates with a number of other tools and frameworks to provide solutions for Knowledge Graph extraction and usage.

### [](#_unstructured_io)Unstructured.io

[Unstructured.io](https://unstructured.io) is a platform and python package for extracting structured data from unstructured documents. It provides a set of tools for parsing, cleaning, and transforming unstructured data into structured formats.

The Neo4j Integration allows you to extract lexical graphs from any supported source.

![unstructured ui neo4j](https://dist.neo4j.com/wp-content/uploads/20250506043320/unstructured-ui-neo4j.png)

Both in the [Platform](https://docs.unstructured.io/ui/destinations/neo4j) and the [open source package](https://docs.unstructured.io/ingestion/destination-connectors/neo4j), you can use the `Neo4jConnectionConfig`, `Neo4jUploaderConfig` to write the extracted lexical data to a Neo4j graph.

```python
Pipeline.from_configs(
    context=ProcessorConfig(),
    indexer_config=LocalIndexerConfig(input_path=os.getenv("LOCAL_FILE_INPUT_DIR")),
    downloader_config=LocalDownloaderConfig(),
    source_connection_config=LocalConnectionConfig(),
...
    chunker_config=ChunkerConfig(chunking_strategy="by_title"),
    embedder_config=EmbedderConfig(embedding_provider="huggingface"),
    destination_connection_config=Neo4jConnectionConfig(
        access_config=Neo4jAccessConfig(password=os.getenv("NEO4J_PASSWORD")),
        username=os.getenv("NEO4J_USERNAME"),
        uri=os.getenv("NEO4J_URI"),
        database=os.getenv("NEO4J_DATABASE"),
    ),
    stager_config=Neo4jUploadStagerConfig(),
    uploader_config=Neo4jUploaderConfig(batch_size=100)
).run()
```

### [](#_langchain)LangChain

The [first implementation for Knowledge Graph extraction](https://medium.com/data-science/building-knowledge-graphs-with-llm-graph-transformer-a91045c49b59) was the [LangChain LLMGraphTransformer](https://python.langchain.com/api_reference/experimental/graph_transformers/langchain_experimental.graph_transformers.llm.LLMGraphTransformer.html).

![1*aCSCXuvrOB90jRQ0mNZtSA](https://miro.medium.com/v2/resize:fit:2000/format:webp/1*aCSCXuvrOB90jRQ0mNZtSA.png)

The usage is pretty straightforward, as documented in [Constructing Knowledge Graphs](https://python.langchain.com/docs/how_to/graph_constructing/). You can use the `convert_to_graph_documents` method to convert a list of LangChain documents into graph documents (nodes and relationships). To guide the extraction process, you can specify a schema with allowed nodes and relationships, as well as properties and additional extraction instructions.

```python
from langchain_experimental.graph_transformers import LLMGraphTransformer

llm_transformer = LLMGraphTransformer(
    llm=llm,
    allowed_nodes=["Person", "Country", "Organization"],
    allowed_relationships=["NATIONALITY", "LOCATED_IN", "WORKED_AT", "SPOUSE"],
    node_properties=["born_year"],
)
graph_documents = llm_transformer.convert_to_graph_documents(documents)[0]

print(f"Nodes: {graph_documents.nodes} Rels: {graph_documents.relationships}")

graph.add_graph_documents(graph_documents, include_source=True)
```

![graph construction4 e41087302ef4c331c2c95b57467f4c62](https://python.langchain.com/assets/images/graph_construction4-e41087302ef4c331c2c95b57467f4c62.png)

### [](#_llamaindex)LlamaIndex

In [LlamaIndex the PropertyGraphIndex](https://www.llamaindex.ai/blog/introducing-the-property-graph-index-a-powerful-new-way-to-build-knowledge-graphs-with-llms) is the main component for managing Knowledge Graphs including building them from unstructured data. For extracting entities and relations, different extractors can be used, for instance the `SchemaLLMPathExtractor`.

```python
from typing import Literal
from llama_index.core.indices.property_graph import SchemaLLMPathExtractor

entities = Literal["Person", "Place", "Thing"]
relations = Literal["PART_OF", "HAS", "IS_A"]
schema = {
    "Person": ["PART_OF", "HAS", "IS_A"],
    "Place": ["PART_OF", "HAS"],
    "Thing": ["IS_A"],
}

kg_extractor = SchemaLLMPathExtractor(
    llm=llm,
    possible_entities=entities,
    possible_relations=relations,
    kg_validation_schema=schema,
    strict=True,  # if false, will allow triplets outside of the schema
    num_workers=4,
    max_triplets_per_chunk=10,
)

graph_store = Neo4jPropertyGraphStore(
    username="neo4j",
    password="<password>",
    url="neo4j+s://xxx.databases.neo4j.io",
)

# creates an index
index = PropertyGraphIndex.from_documents(
    documents,
    property_graph_store=graph_store,
    # optional, neo4j also supports vectors directly
    vector_store=vector_store,
    embed_kg_nodes=True,
)
```

The Neo4j-PropertyGraphIndex can be used for GraphRAG as a retriever, e.g. via VectorContextRetriever, LLMSynonymRetriever, TextToCypherRetriever, CypherTemplateRetriever.

-   [Property Graph Index Documentation](https://docs.llamaindex.ai/en/stable/module_guides/indexing/lpg_index_guide/)
    
-   [Neo4j Property Graph Index](https://docs.llamaindex.ai/en/stable/examples/property_graph/property_graph_neo4j/)
    
-   [Property Graph Construction in Neo4j](https://docs.llamaindex.ai/en/stable/examples/property_graph/property_graph_advanced/)
    

## [](#_research)Research

-   [GLiNER](https://graphrag.com/appendices/research/2311.08526/), [GLiNER Neo4j](https://bluetickconsultants.medium.com/dual-approaches-to-building-knowledge-graphs-traditional-techniques-or-llms-400fee0f5ac9), [LangChain GlinerGraphTransformer](https://python.langchain.com/api_reference/experimental/graph_transformers/langchain_experimental.graph_transformers.gliner.GlinerGraphTransformer.html)
    
-   [Entity Linking and Relationship Extraction With Relik](https://neo4j.com/blog/developer/entity-linking-relationship-extraction-relik-llamaindex/)
    
-   [SciPhi Triplex](https://www.sciphi.ai/blog/triplex), [Triplex Neo4j](https://www.sciphi.ai/blog/graphrag)
    
-   [KGGen: Extracting Knowledge Graphs from Plain Text with Language Models](https://arxiv.org/abs/2502.09956)
    

## [](#_blog_posts)Blog Posts

-   [From Unstructured Text to Interactive Knowledge Graphs using LLMs (2025)](https://robert-mcdermott.medium.com/from-unstructured-text-to-interactive-knowledge-graphs-using-llms-dd02a1f71cd6)
    
-   [GraphRAG in Action: From Commercial Contracts to a Dynamic Q&A Agent (2024)](https://neo4j.com/blog/developer/graphrag-in-action/)
    
-   [Using LlamaParse for Knowledge Graph Creation from Documents (2024)](https://medium.com/neo4j/using-llamaparse-for-knowledge-graph-creation-from-documents-3bd1e1849754)
    
-   [Automate Building Knowledge Graphs from Scientific Abstracts with LLMs and Neo4j (2024)](https://blog.gopenai.com/automate-building-knowledge-graphs-from-scientific-abstracts-with-llms-and-neo4j-2-ways-b9d5152a12da)
    
-   [Enhancing the Accuracy of RAG Applications with Knowledge Graphs (2024)](https://medium.com/neo4j/enhancing-the-accuracy-of-rag-applications-with-knowledge-graphs-ad5e2ffab663)
    
-   [Document Knowledge Graph with Unstructured.io (2024)](https://neo4j.com/blog/developer/document-knowledge-graph-unstructuredio/)
    
-   [Constructing Knowledge Graphs from Text using OpenAI Functions (2023)](https://bratanic-tomaz.medium.com/constructing-knowledge-graphs-from-text-using-openai-functions-096a6d010c17)
    
-   [Text to Knowledge Graph Information Extraction Pipeline (2022)](https://neo4j.com/blog/genai/text-to-knowledge-graph-information-extraction-pipeline/)

---
_Imported from <https://neo4j.com/developer/genai-ecosystem/importing-graph-from-unstructured-data/>_


---

## Summary
Modern large language models enable the automated extraction of entities and relationships from unstructured data to build knowledge graphs. These graphs can be stored in Neo4j and utilized for applications such as GraphRAG to improve the contextual accuracy of generative AI responses.

## Themes
- Knowledge Graph Construction
- LLM-based Entity Extraction
- GraphRAG Pipeline Integration
- Unstructured Data Processing

## Key takeaways
- LLMs can extract and de-duplicate entities and relationships from unstructured text, images, and audio to populate graph databases.
- The Neo4j Python GraphRAG package provides a comprehensive pipeline for document processing and graph storage.
- Tools like the LLM Knowledge Graph Builder and LangChain LLMGraphTransformer automate the conversion of documents into graph structures.
- Knowledge graphs improve RAG applications by providing relevant contextual information to augment LLM responses.
- Ecosystem integrations with Unstructured.io, LangChain, and LlamaIndex facilitate standardized workflows for building knowledge graphs from various document formats.

## Insights
- Modern LLMs have shifted the burden of knowledge graph creation from manual NLP model fine-tuning to prompt-based schema definition.
- The construction of a robust knowledge graph requires a multi-stage pipeline including text splitting, entity resolution, and schema grounding to ensure navigability.
- Integration with graph databases like Neo4j allows for the combination of unstructured extracted data with existing structured datasets for enhanced RAG performance.

## Open questions / gaps
- What are the specific performance trade-offs between different LLM providers when performing large-scale entity extraction?
- How does the choice of graph schema impact the long-term maintainability of a knowledge graph as source data evolves?

## Concepts in this document
- **Neo4j** _(entity)_
  Graph database platform that serves as the storage and retrieval foundation for knowledge graphs and hybrid search.
- **GraphRAG** _(concept)_
  Retrieval-augmented generation pattern that uses knowledge graphs to provide contextual information for LLM responses.
- **Knowledge Graph** _(concept)_
  Structured representation of entities and relationships extracted from unstructured data for contextual retrieval.
- **Hybrid Search** _(concept)_
  Multi-signal retrieval combining lexical, semantic, and structural search to improve result quality and coverage.
- **Graph Database** _(entity)_
  The technology being investigated as a foundational architecture for personal knowledge management.
- **Knowledge Management** _(tag)_
  Broad domain of organizing, storing, and retrieving information for personal or organizational use.
- **Memory system** _(concept)_
  Concept describing graph databases as a memory system for work.
- **Neo4j Graph Data Science** _(entity)_
  Neo4j library providing algorithms like FastRP for converting graph topology into searchable vector embeddings.
- **Vector indexes** _(concept)_
  High-dimensional vector representations enabling semantic similarity in Neo4j.
- **Retrieval-Augmented Generation** _(tag)_
  Domain combining information retrieval with generative AI to provide contextually grounded LLM responses.
- **Apache Lucene** _(entity)_
  Indexing and search library that powers Neo4j vector indexes.
- **memory.wiki** _(entity)_
  A knowledge management platform providing REST APIs, CLI tools, and MCP server integration.

## Concept relations (within this doc's concepts)
- **Knowledge Management** contextualizes exploration of **Graph Database**
- **GraphRAG** uses to augment **Knowledge Graph**
- **GraphRAG** utilizes **Knowledge Graph**
- **Neo4j** is a **Graph Database**
- **Neo4j** supports **Hybrid Search**
- **Neo4j** implements **Vector indexes**
- **memory.wiki** hosts documents on **Neo4j**
- **Vector indexes** powered by **Apache Lucene**
- **Hybrid Search** improves results for **Retrieval-Augmented Generation**
- **GraphRAG** shares concept **Knowledge Graph**
- **Graph Database** implemented by **Neo4j**
- **Memory system** implemented via **Graph Database**
- **Neo4j** is implementation of **Graph Database**
- **Neo4j** supports advanced **Hybrid Search**
- **Neo4j** stores and manages **Knowledge Graph**
- **Neo4j** is type of **Graph Database**
- **Neo4j** features **Vector indexes**
- **memory.wiki** hosts documentation for **Neo4j**
- **GraphRAG** utilizes structure of **Knowledge Graph**
- **Neo4j** supports implementation of **Hybrid Search**

## Bundles containing this document
- [Neo4j and GraphRAG implementation](https://memory.wiki/b/bMK8g0yP)

_Hub canonical:_ https://memory.wiki/hub/pratofeito
_Concept digest:_ https://memory.wiki/raw/hub/pratofeito?digest=1&compact=1
