On this article, you’ll learn to add a light-weight temporal reasoning layer to a Graph-RAG system in order that it could possibly distinguish contemporary information from stale ones.
Subjects we’ll cowl embody:
- Methods to lengthen customary subject-predicate-object triples into time-stamped quadruples saved in a easy temporal graph.
- Methods to calculate recency weights with exponential decay and use them to rank conflicting information as of a given question date.
- Methods to tune the half-life parameter and combine the temporal graph right into a deterministic 3-tiered Graph-RAG retrieval pipeline.
Introduction and Motivation
In a earlier article, Constructing a Deterministic 3-Tiered Graph-RAG System, we addressed the problem of dealing with conflicting data in RAG (Retrieval-Augmented Era) architectures. Particularly, we constructed a hierarchy to handle conflicting data by giving “contemporary” information high precedence over less-fresh ones.
All through that journey, a key query arose: how does our graph-based RAG system know precisely what’s contemporary? Customary data graphs deal with information as “timeless,” context-independent SPO (subject-predicate-object) triples, reminiscent of (Firm, HAS_CEO, Alice). This method doesn’t fairly match the true world we stay in, the place issues are messy and alter continually: what if Alice switched jobs and is not CEO? Feeding these information to an LLM with out temporal context is the right recipe for hallucinations, leading to complicated or factually incorrect responses.
To sort out this problem, this text reveals the key steps to construct a devoted, light-weight temporal reasoning engine for Graph-RAG via a number of easy Python features. We additionally focus on its integration with the 3-tiered Graph-RAG system constructed beforehand. The important thing thought consists of upgrading customary triples into time-stamped quadruples and calculating recency weights that point out levels of “truth freshness.”
Prepared? Let’s go!
A New, Temporal Journey, Step by Step
Step one is to increase customary SPO triples into “temporal quads,” the place the fourth dimension introduces time, concretely a timestamp: (Topic, Predicate, Object, Timestamp).
The next Python class is outlined to carry our new, prolonged data, additionally referred to as a temporal graph. In case you are working in a pocket book atmosphere, merely paste this code into your first code cell:
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 |
import datetime import math
class TemporalGraph: def __init__(self): # Knowledge might be saved on this data base as: Topic -> Predicate -> Record of (Object, Date) self.knowledge_base = {}
def add_fact(self, topic, predicate, obj, date_string): “”“Provides a time-stamped truth to the graph.”“” # Changing string to a comparable date object fact_date = datetime.datetime.strptime(date_string, “%Y-%m-%d”).date()
if topic not in self.knowledge_base: self.knowledge_base[subject] = {} if predicate not in self.knowledge_base[subject]: self.knowledge_base[subject][predicate] = []
self.knowledge_base[subject][predicate].append((obj, fact_date)) print(f“Added: {topic} {predicate} {obj} (as of {date_string})”)
# Initializing our graph tg = TemporalGraph() |
Now it’s time to populate our newly created temporal graph, tg, following a real-world state of affairs the place information change at mild velocity … effectively, possibly not that quick, however nonetheless quickly! If we have been monitoring the management roles in a tech firm throughout a chaotic week stuffed with modifications, we might have one thing like:
|
# A timeline of shifting information tg.add_fact(“TechCorp”, “HAS_CEO”, “Alice”, “2021-01-15”) tg.add_fact(“TechCorp”, “HAS_CEO”, “Bob”, “2023-11-17”) tg.add_fact(“TechCorp”, “HAS_CEO”, “Charlie”, “2023-11-19”) tg.add_fact(“TechCorp”, “HAS_CEO”, “Bob”, “2023-11-21”) # Bob got here again!
# Including additionally a static truth for a little bit of distinction tg.add_fact(“TechCorp”, “FOUNDED_IN”, “San Francisco”, “2010-05-01”) |
Output:
|
Added: TechCorp HAS_CEO Alice (as of 2021–01–15) Added: TechCorp HAS_CEO Bob (as of 2023–11–17) Added: TechCorp HAS_CEO Charlie (as of 2023–11–19) Added: TechCorp HAS_CEO Bob (as of 2023–11–21) Added: TechCorp FOUNDED_IN San Francisco (as of 2010–05–01) |
Keep in mind that in a regular RAG system, a search like “Who acts because the CEO of TechCorp?” would doubtless retrieve Alice, Bob, and Charlie, unexpectedly! Thus, we’d like a mechanism to assign truthfulness weights to information, and it’s less complicated than you would possibly assume.
Calculating recency weights is the important thing to resolving attainable conflicts mathematically. We simply need a mechanism that claims: “hey, this truth is newer than that one, so it’s extra prone to represent as we speak’s fact.” A wise method to do that is predicated on exponential decay, which consists of assigning a half-life time window to information. As an example, if the half-life is ready to at least one 12 months (three hundred and sixty five days), then a truth that’s one 12 months previous will carry a weight of 0.5. In the meantime, a truth asserted as we speak would carry a weight of 1.0.
These two features are designed to introduce the aforementioned weight scoring logic to our temporal graph:
|
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 |
def calculate_recency_weight(fact_date, query_date, half_life_days=365): “”“ Calculates a rating between 0 and 1 based mostly on how previous the very fact is. Utilizing exponential decay: weight = (0.5) ^ (age_in_days / half_life) ““” age_in_days = (query_date – fact_date).days
# If the very fact is from the long run relative to our question, it’s capped at 1.0 if age_in_days < 0: return 1.0
weight = 0.5 ** (age_in_days / half_life_days) return spherical(weight, 4)
def query_temporal_graph(graph, topic, predicate, as_of_date_str, half_life_days=365): “”“Queries the graph and ranks solutions by their temporal weight.”“” query_date = datetime.datetime.strptime(as_of_date_str, “%Y-%m-%d”).date()
strive: information = graph.knowledge_base[subject][predicate] besides KeyError: return f“No data discovered for {topic} -> {predicate}”
scored_results = [] for obj, fact_date in information: # We solely contemplate information that occurred ON or BEFORE our question date if fact_date <= query_date: weight = calculate_recency_weight(fact_date, query_date, half_life_days) scored_results.append({ “reply”: obj, “date”: fact_date.strftime(“%Y-%m-%d”), “weight”: weight })
# Sorting by weight (highest/freshest first) scored_results.type(key=lambda x: x[‘weight’], reverse=True) return scored_results |
Lastly, we’re able to see all of it in motion. We are going to end by displaying an instance that queries our graph. Temporal reasoning acts as a sort of “time journey” at execution time: if we added code to persist our information after which requested who the CEO was a number of days later, the mechanism we applied would merely modify its weights on the fly:
|
print(“— Question 1: Who’s the CEO as of Nov 18, 2023? —“) results_past = query_temporal_graph(tg, “TechCorp”, “HAS_CEO”, “2023-11-18”) for res in results_past: print(f“Candidate: {res[‘answer’]} | Truth Date: {res[‘date’]} | Confidence Weight: {res[‘weight’]}”)
print(“n— Question 2: Who’s the CEO as of Dec 01, 2023? —“) results_present = query_temporal_graph(tg, “TechCorp”, “HAS_CEO”, “2023-12-01”) for res in results_present: print(f“Candidate: {res[‘answer’]} | Truth Date: {res[‘date’]} | Confidence Weight: {res[‘weight’]}”) |
Outcomes:
|
—– Question 1: Who is the CEO as of Nov 18, 2023? —– Candidate: Bob | Truth Date: 2023–11–17 | Confidence Weight: 0.9981 Candidate: Alice | Truth Date: 2021–01–15 | Confidence Weight: 0.1396
—– Question 2: Who is the CEO as of Dec 01, 2023? —– Candidate: Bob | Truth Date: 2023–11–21 | Confidence Weight: 0.9812 Candidate: Charlie | Truth Date: 2023–11–19 | Confidence Weight: 0.9775 Candidate: Bob | Truth Date: 2023–11–17 | Confidence Weight: 0.9738 Candidate: Alice | Truth Date: 2021–01–15 | Confidence Weight: 0.1362 |
As one would possibly anticipate, operating the primary question provides us Bob as the highest reply with an virtually full weight: Charlie doesn’t even seem, as he hadn’t been appointed at that time! In the meantime, operating the second question reveals a caveat: maybe the 365-day half-life window is just too lengthy, because it takes a complete 12 months for information to lose 50% of their relevance. Thus, in a frenetic week stuffed with organizational modifications, we are able to see that although Bob is once more the highest reply, he’s very intently adopted by Charlie and even by Bob’s personal prior appointment. The fast repair consists of adjusting the half_life_days parameter, as an example, by altering it from 365 to 7. Strive it your self and revel in the brand new outcomes!
Wrapping Up
Now that now we have constructed this mechanism to take temporal graph data into consideration, how might it’s built-in into the deterministic 3-tiered structure constructed within the earlier, associated article? When the consumer sends a immediate to the LLM within the RAG system, you’ll need a retriever that not solely fetches texts: as an alternative, it ought to run the question towards the temporal graph, type information by their confidence weight, and cross solely the top-weighted one (or, at most, a small ranked listing) into the immediate’s context. This has the potential to take away the LLM’s must guess which truth is probably the most present one: that situation is sorted out even earlier than the ultimate immediate reaches the mannequin.

