Roman Klimenko
DA

CluedIn Python SDK 1.0.0

I've released cluedin 1.0.0, a Python SDK for the CluedIn data management platform.

The library is available on PyPI and GitHub. It's open source and free to use.

It focuses on authentication and GraphQL API support. I plan to add more features in later releases.

Installation

Install the package with:

pip install cluedin

Usage

Authentication

To get a CluedIn access token, create a context object and pass it to cluedin.load_token_into_context(context):

import cluedin

context = {
    "protocol": "http", # if you skip this parameter, it will fall back to `https`
    "domain": "cluedin.local",
    "organization": "foobar",
    "user": "admin@foobar.com",
    "password": "Foobar23!"
}

cluedin.load_token_into_context(context)

print(context['access_token'])

GraphQL

To run a GraphQL request, pass the context object to cluedin.gql.gql(context, query, variables):

query = """
  query searchEntities($cursor: PagingCursor, $query: String, $pageSize: Int) {
    search(
      query: $query
      sort: FIELDS
      cursor: $cursor
      pageSize: $pageSize
      sortFields: {field: "id", direction: ASCENDING}
    ) {
      totalResults
      cursor
      entries {
        id
        name
        entityType
        properties
      }
    }
  }
"""

variables = {
    "query": "entityType:/Infrastructure/User",
    "pageSize": 1
}

response = cluedin.gql.gql(context, query, variables)

To retrieve every page of results, use the cluedin.gql.entries(context, query, variables) generator:

import numpy as np
import pandas as pd

query = """
  query searchEntities($cursor: PagingCursor, $query: String, $pageSize: Int) {
    search(
      query: $query
      sort: FIELDS
      cursor: $cursor
      pageSize: $pageSize
      sortFields: {field: "id", direction: ASCENDING}
    ) {
      totalResults
      cursor
      entries {
        id
        name
        entityType
        properties
      }
    }
  }
"""

variables = {
    "query": "*",
    "pageSize": 10000
}

entries = np.array([x for x in cluedin.gql.entries(context, query, variables)])

df = pd.DataFrame(entries.tolist(), columns=list(entries[0].keys()))

If you run into trouble or have an idea for the library, let me know.

Release notes

Authentication

  • cluedin.auth.get_token_response(context): gets an access token response.
  • cluedin.load_token_into_context(context): loads a JWT access token into a context object.

GraphQL

  • cluedin.gql.gql(context, query, variables): runs a GraphQL request.
  • cluedin.gql.entries(context, query, variables): returns paginated GraphQL results as a generator.

URLs

Utilities

  • cluedin.utils.load(filename): loads a JSON file into an object.
  • cluedin.utils.save(obj, filename, sort_keys=True): saves an object to a JSON file.

PyPI package: https://pypi.org/project/cluedin/1.0.0/