The dilemma and cost of making data accessible and usable

Introduction

Happy New Year everyone.

I was inspired by a host of innovative, exciting speakers following my attendance and Chairing of Day 1 at the Public Spectrum DigitalGov conference in Canberra late last year. I went home with a lot of appetising food for thought around the question of making data accessible and usable, and the dilemma and cost around making that work. How can we get there?!

In the research community we have been dabbling with how to do this in a standardised way with lot of different types of data, for some time. So, it seems there is light at the end of the tunnel, we just have to keep dabbling!

DigitalGov, Data and Cloud

At the conference, discussions around Cloud First seem to have progressed over the last few years. Agencies have now become quite mature in their cloud journeys and have now migrated most of their apps into the cloud where they can be scaled in a hybrid cloud scenario. But now we found disparate data across these apps and cloud services that need to be linked up, increasing interconnectivity, and therefore enabling data usability and shareability.

I enjoyed the opportunity to Chair the sessions on Day 1, which exposed the audience to innovative thought leadership around topics, such as; “Realising that digital transformation will never end but is always unfolding” (Beth Killoran), and “Technology solutions: Maximising the true potential of your data assets” (Ben Henshall).

I heard the phrase data virtualisation for the first time at the conference. Data virtualisation really refers to data caching or a data services layer. This has been routinely used by the geospatial community for well over 10 years when they discovered the value of retaining data at source (where it can be updated regularly), but it can also be delivered via web-service to enable user’s ‘live’ access.  Thinking about data infrastructure in this way will enable data usability in the future. Watch this space!

Data Access Challenges

The conference, in particular got me thinking about how we can improve ways to make data more accessible and usable. However, overcoming the dilemmas and costs around that are major barriers. But if we can make progress in this area, it will transform the way we do things

If we get back to basics – data accessibility is about consolidating data to one easy-to-access location or platform to view, for citizens, a company workforce, or you, the researcher (e.g. data democratization).

In a world where data is fully accessible, there are no gatekeepers or barriers preventing access to others data whenever and wherever you want. But with this opportunity comes great responsibility to ensure we are using data in ethical and safe ways.

Data Management

What makes a real difference to researchers on the ground and can be a useful lesson for everyone in business now is good data management. It is the key to ensuring data is easy to use in the future making it discoverable and better documented.

The Research Data Alliance (RDA) is the go-to authority on research data management globally, and provides a single point of truth for all research data. It has established the goal posts and framework so data can be openly shared. The RDA provides data consistency and standards so that researchers and user infrastructures have a single way to interact with the data.

The RDA is paving the way for the future by encouraging the accessibility and useability of digital data. Innovative thinkers outside of research can learn from the insights achieved at RDA.

Looking to the governance framework provided by the RDA, we can seek to adopt some best practices here in Australia to enable research and innovation. If we can streamline the process then the research community can focus on what it does best – undertake data analysis and research – without having to factor in the always time-consuming nature of data management, creation of bespoke data or contend with the sometimes-finite capability of company data storage infrastructure.

Data Capital

However, we can’t underestimate the resources (people and tools) needed to make data accessible and interoperable. It’s often said, “our people are our capital” but when it comes to the digital data they hold, companies are realising that their data is one of their top assets. After all, they have invested in its initial acquisition, processing and storage.

Much of the data we generate is underused, which means Australia’s data is ripe for the picking! It’s a fact – data stored by your organisation, needs to be accessible right now, but with broader data sharing, we must consider the right to privacy, security and intellectual property (IP) of data owners (Productivity Commission, 2017).

There will probably always be a trade-off between data availability and enabling future accessibility. The question is – how do we prioritise data access?

FAIR Data

We need to improve how we can provide data as a service to the users of the data, like researchers, so they can focus on analysis and delivery. Removing any unnecessary roadblocks, that threaten to limit or hinder access.

A major step forward to accessing reliable open data sources for researchers, is having Findable, Accessible, Interoperable and Reusable (FAIR) data. This system provides a consistent approach, with institutions and repositories required to meet a core trust standard (RDA, 2019).

The lower privacy and security limitations of Earth, Space, and Environmental Science data, allowed journals in these fields to meet the standard first. This has seen 95% of data cited, now open and FAIR.

Data Transformation

Without gatekeepers restricting access, the only real barriers to the access and use of digital data is data transformation. Re-formatting, re-structuring, or modifying (adding or deleting) values of data is an additional layer of processing, done to suit the end-user.

Let’s weigh up the costs. With the falling costs (per record) of digital data storage and the massive benefits and opportunities available from openly accessible data, costs should be less of an impediment these days.

We could look at a tiered approach to the cost of making data open and accessible (Productivity Commission Report on Data Availability and Use, 2017). Some readily available data, with low privacy restrictions can probably be made available quickly. Whereas, lower quality or specialised data may take longer to transform into a usable format, therefore producing higher cost and taking longer to make it accessible.

Final word

Professor Michael Blumstein, Associate Dean Research at UTS said that we need to focus on all the new data sources we are creating first, like Internet of Things (IoT), etc. I agree, because we don’t want to create more legacy debt in our data sources. But our older Data is an asset and has value in helping us prepare for the future, so we still need to devote resources (time, budget and people) to manage this data.

I would recommend that when we think about making data accessible and usable, we need to ask the question – what do we want to know and where are the data sources that can help us? Then we can prioritise the data that you need to remediate – in such a way that is sustainable and future proofed, and enables new data to be added to the data source once collected.

It is in the best interest of companies to increase the accessibility and useability of their digital data, investing in standardised approach by using the example set by the RDA and tools for data transformation. If it’s done right, it better enables research and innovation today, but also ensures its readiness for future opportunities.

References

Productivity Commission 2017, Data Availability and Use, Report No. 82, Canberra.

RDA Research Data Alliance, 2019. Enabling FAIR Data in the Earth, Space, and Environmental Sciences (https://www.youtube.com/watch?v=K7jLPfgFXXU&list=PL77Dq-HHs58a9GRarka68-1OR5MAe8SPk&index=1 ).

Leave a Reply