Exercise 3.7. Design a Company knowledge graph to support a business intelligence dashboard. The dashboard is to aggregate information from multiple sources about a company to get better insight into its business, customers, competitors, subsidiaries or parent organlzations. Assume that the two sources are Wikidata and Security and Exchange Commission Filings. Begin the process by reviewing the RDF schema for Company (identifier Q783794) in Wikidata. Reuse as much of this schema as necessary, and extend it as you see fit. Document your choices. For SEC filings, use the FinReflectKG dataset. Extend the schema you had extracted from Wikidata in the previous step to handle any new information that appears in FinReflectKG. Examining Company Schema in WikidataWe examine the company schema in Wikidata by posing the following SPARQL query in the Wikidata query interface.
The query returns more than 250 properties. Many of these properties are not relevant to the current task. Let us pick the following properties to work with.
Examining FinReflectKG SchemaFinReflectKG is a financial knowledge graph dataset extracted from S & P 500 companies' SEC 10-K filings spanning 2014-2024, containing 17.51 million normalized triplets with full textual context. The following discussion is based on the schema description included with the dataset. The core elements of the FinReflectKG triple are shown below.
The above triples could be straightforwardly mapped into RDF triples, but before jumping into that conclusion, let us examine the contents of the triples more closely. The entity_type and target_type in this dataset are quite diverse, ranging from organizations, people, events, abstract concepts, raw materials, policies, etc. We will, therefore, need to filter the dataset so that it best adds to the design we are creating here. The dataset also has an extensive set of relationship types. In principle, all of the relationships could be in a company knowledge graph, but to maintain a reasonable scope for this exercise, we choose the following relationships.
We have chosen the above relationships to closely correspond to several of the Wikidata properties introduced earlier, although the mapping is not always exact. For example, produces aligns directly with product or material produced (P1056), while parent_of and subsidiary_of correspond to organizational hierarchy relationships such as child organization or unit (P355) and part of (P361). In contrast, invests_in and has_stake_in have only approximate equivalents in owned by (P127), since they distinguish between investment, partial ownership, and corporate control. Finally, supply has no close counterpart among the Wikidata properties considered here, reflecting FinReflectKG's emphasis on modeling business relationships such as supply chains that are not explicitly represented in the selected Wikidata schema. Company Knowledge Graph schemaFor this exercise, we will use the Wikidata schema as the foundation. We will use the Wikidata identifiers for the companies. We will import only those triples from FinReflectKG that satisfiy the Wikidata property constraints. The key entity types in our schema will be Entity which can be either and Company or a Person, Industry, and Offering (a generalized label for prodcut or service) . As invests_in and has_stake_in have only approximate equivalents in owned by (P127), we will add them as additional relationships. We will add the supply relationship to our schema.
|