Showing posts with label data. Show all posts
Showing posts with label data. Show all posts

Thursday, November 15, 2018

Data and Software Carpentries

What is Data Science?

Well, simply put, it is the act of collecting new and insightful ideas and methods from data sets. The modern world revolves around data, and data is an important factor in any business or organization. You have your finances, your social media interaction, your production, your employee and partner information - the list goes on. No matter where you are in life, you’re going to have to deal with data at some point. Data Science deals specifically with digital data, which is based on various systems of code that are interpreted, processed and altered by computers.

The Data and Software Carpentries are programs that are looking to teach people how to conveniently and effectively manage the data that they gather in their businesses, organizations and even research opportunities at school. The benefits are that research becomes less painful, and the techniques learned through the data carpentries lead to more research being done more effectively.

According to Phillip Doehle, the program coordinator in the Oklahoma State University High Performance Computing Center, the tagline of the Carpentries is to “help researchers get their work done with less pain and more convenience.”

Software and Data Carpentry are not the same thing, though there is much overlapping between the two and many of the same tools are used. The difference is that Data Carpentry deals with a method of organizing data in an easy and understandable way so that interpretation and analysis are more simple and convenient. A huge factor that plays into Data Carpentry is the Data Lifecycle, which consists of data collection, data organization/cleaning, data analysis and a positive production trend.

Alternatively, Software Carpentry focuses on programming skills and techniques such as software development. Coding is a big tool within this program, which is used to create a reproducible method. According to Evan Linde, a research cyberinfrastructure analyst from the OSU High Performance Computing Center, Software Carpentry can give you the introductory knowledge needed to get you out of a tight situation. For example, a graduate student struggling with a data problem can learn how to write a specific bit of code needed to link multiple processes together and solve that problem.

Now you might be thinking, “I have no experience with stuff like this, and I don’t think I’ll be able to keep up with it!” According to the official Data Carpentry website, their initial target audience is “learners with little to no computational experience.” Even if there is a bit of a learning curve, the experience and benefits gained are worth the time investment.

Software and Data Carpentries are starting points for coding and data management. The idea of the Carpentries is “Reproducible Data.” What this means is that the data can be duplicated through different means. When a data set is duplicated through different methods, then error risks are reduced, results are validated and it shows that the data can be adaptive to multiple different conditions.

While the techniques learned in the Carpentries can be very useful and informative, they don’t necessarily make you ready to immediately go out and work at a research facility. The Carpentries have a larger emphasis on the academics of research rather than an industrial emphasis. Even so, they provide a beneficial and explanatory introduction to the world of Data Science and Data Analysis. The programs are beneficial for graduate students who are getting involved in research or are interested in concepts of data science. Faculty members have told me that the knowledge they gained from the Data and Software Carpentries helped them maximize the time that they spent as a graduate student when they were working with research because their data was organized and reproducible.

The Carpentries are an informative and practical method for somebody who is looking to delve into a career based around data science, as well as students who are involved in research opportunities. They provide a foundation for understanding how to organize and interpret data, as well as how to make that data applicable within various situations. They’re an opportunity that many don’t know about, and they can potentially give somebody an extra push in understanding their career or research choice.

Posted by Braiden Ellis, student staff member, Research and Learning Services

Thursday, October 25, 2018

Data Culture in the Library

If I asked somebody what a data culture is, they would probably look at me with a disinterested, glazed expression, followed closely by a mumbled, “I don’t know.”

And they may not really care either.

I feel like most people tend to see the word “data” as a very scary and intimidating word, and this prevents them from learning about just how useful and informative data can really be. Most people also tend to think that data is strictly quantitative information, such as numerical statistics. However, data can be text as well (such as emails, interview notes and media reports), which is qualitative data.

When analyzed, data provides ideas and solutions to various problems and situations. This ties into data culture and the benefits it can bring to a company or organization.

A data culture is an organization that uses data to make decisions to improve themselves. Some business leaders may prefer to make decisions based on gut instinct, ignoring data because their way just “makes more sense” (at least to them it does). A data-driven organization follows the path that makes the most sense in correlation with the data, even if it doesn’t necessarily make sense from a conventional standpoint. This has proven to be far more strategic and efficient. Erik Brynjolfsson et al. al. from MIT’s Sloan School of Management conducted a study that showed data-driven organizations to have a five to six percent higher output and productivity than other systems in which organizations tend to go with what they “feel” is the best path.

The Data Culture Project provides a great method for learning how data can be useful and how to connect with others through detailed analysis. The Data Process described below is a step-by-step system from the Data Culture Project that details how to bring about a more data-oriented workplace. The OSU Library is piloting the Data Culture Project this semester and plans to offer it in the future for those who are interested.

The Data Process:

#1: Warmups
Anytime you get a group together for a team-building exercise, warmups are a great way to get things started. Quick and creative activities that help to establish a feeling of comfort and informality allow participants to be invested and confident, leading to more engagement, which leads to more results in the long run.

#2: Asking Questions
This is the first step in analyzing data. When we look at a data set, there are plenty of questions to be asked, and there are plenty of questions to be answered. Colleagues are divided into small groups to analyze a data set and develop insightful observations and questions. Now, this exercise isn’t particularly about answering any questions that are asked, but it is instead about thinking creatively to generate stimulating questions.

#3: Gathering Data
This step goes hand-in-hand with the previous step. Data is difficult to interpret sometimes, even more so when the data is wrong or unorganized. You might need to combine multiple data sets as well, because a single data set may not be able to answer every question.

#4: Analyzing Data
Oh no, here comes the hard stuff!
Not really, but this is how many people feel when it comes to analyzing data. There are many different ways to analyze data because there are many different types of data. If you’re having a hard time trying to figure out a data set, maybe you just aren’t using the right method for you!  Analyzing data with colleagues leads to a combination of different ideas to maintain positive trends.

#5: Telling Your Story
This is the fun part of the process (as if you weren’t having fun before). This allows you to be creative while explaining a data set to colleagues. Various types of displays such as sketches and word webs communicate the context of the data through a narrative. This also allows abstract ideas and numbers to be represented by a concrete visual display.

#6: Try It Out
The final step is simple: just try it out! This answers many questions you may have about your developed story and the data backing it. Is the data understandable? Does the story make sense? Is it compelling? The idea of narrating your data is to influence your audience, informing and propelling them into action to improve the data.

A data-driven culture isn’t something that immediately happens, and it definitely isn’t something that has immediate results. It requires cooperation and understanding among colleagues, and the entire organization needs access to data in order to practice analyzing and interpreting it. In addition, data literacy must be encouraged so that each member of the organization can effectively contribute to improving and developing a collective network to help make important decisions.

Posted by Braiden Ellis, student staff member, Research and Learning Services

Monday, July 23, 2018

Learning about Agriculture Data and Databases

In June, I attended a conference called Driving Innovation through Data in Agriculture. It was an enjoyable and interesting experience, combining researchers and librarians onsite at the National Agricultural Library (NAL). There were several things that I brought back with me for further discussion. One of them was that I learned quite a bit about the different agriculture specific databases that the USDA manages. They are all public and provide different specialty products and can be accessed through the National Agricultural Library website. The NAL is responsible for building an integrated portfolio of knowledge products that can be used by the general public, researchers and those working in agriculture.

The Long-Term Agroecosystem Research Data Overview (LTAR) database publishes the data and information generated through the Long-Term Agroecosystem Research (LTAR) Network, a partnership of 18 research sites across the United States. They focus on building knowledge that will make agriculture sustainable and examining the environmental impacts of farming practices. The data generated is “Common Observatory Data” or CORe observational data and research data from management studies with crops and livestock. It compares current practices to experimental practices to determine the effects of the changes. Available data types from LTAR include site photographs, weather and aerial images for the research stations as well as data sets and reports published by LTAR. It has an average of 50 years of historical data. LTAR maintains the data to be sure that it is FAIR: Findable, Accessible, Interoperable and Re-usable. It is hosted at the Ag Data Commons repository and the Geospatial Platform. The Ag Data Commons is the central registry for USDA data and provides links to any data that is housed outside of the repository, including data that is first uploaded to agency sites and provides citable DOI’s for each dataset. The Geospatial Platform provides curation for GIS data but is cross referenced and linked with both Ag Data Commons and Data.gov, the government’s one stop shop for all government sponsored data.

Some of the other specialized USDA data bases include the Life Cycle Assessment Commons, USDA Food Composition Database, the i5 Workspace@NAL, which is a community specializing in arthropod genomes and research and Dr. Duke’s Phytochemical and Ethnobotanical Databases. Less specialized databases include the NAL catalog which covers the entire NAL collection, AGRICOLA, PubAg and NAL Digital Collections, which includes all the full text digitized materials of NAL.

I hope researchers, faculty and students will explore some of these. I would be interested to hear about any projects that utilize them. We here in the library are discussing the use of these specialized databases versus the more popular, general use ones. Feel free to post any examples and please, contact me or any of our friendly librarians if you need more information.

post by Kay Bjornen, Research Data Initiatives Librarian

Monday, May 21, 2018

Introducing Kay Bjornen: The New Research Data Initiatives Librarian

A Library Future Series post from February 13, 2017, highlighted the work done to establish data management services on the Oklahoma State campus. The post explained the survey and data collection process that was used to identify needs and target services that were most needed by researchers here at OSU. Since then there have been several important developments.

One of them is that I was hired to serve as a full-time data services librarian. My name is Kay Bjornen, and I joined the library as a Research Data Initiatives Librarian in February. I bring a Ph.D. in chemistry and a long career in research and data management with me. Since research support and data management are my sole responsibilities (at least for now), I have had the opportunity to knock on doors and introduce myself to many of the departments on campus that carry on data intensive research. I want to make sure that faculty and students know I am here, and I am anxious to assist with any needs they may have. If I haven’t made it to your department yet, feel free to reach out. Or just hold tight, I plan to get there as soon as possible.

One of the questions I have been asked is "Why is the library doing this?" The answer is research is changing and so are libraries. Researchers have access to a wealth of information resources and need less assistance to conduct their literature searches, chase down interlibrary loan items and answer reference questions. Information is increasingly online and digital so librarians are comfortable with electronic information sources of all kinds. It is also true that in the era of "Big Data," there are more and more challenges in managing research outputs. These include describing, storing, securing, organizing and sharing data. The library is a natural partner in these things. That’s why we are working to get the word out about the services we offer.

I am excited to be here at OSU and anxious to contribute to the research community. Do you have any questions about how the library can help you with your research? Feel free to post them here or check out the research data services we offer on the library website. My contact information is there and so is a consultation form you can fill out with the information I need to help you with your research data dilemma.

Post by Kay Bjornen, Research Data Initiatives Librarian

Monday, December 26, 2016

Rooms of Purpose

It is well-known by now that the OSU Library has seen some changes in recent times. With the new equipment in the Creative Studios and the renovation of the 1st floor west wing, our aged building continues to be a hotspot for students while cultivating a venue for creation and innovation. Some of its most useful rooms are sometimes overlooked, so the following are several spots that may not receive as much attention as they deserve.


Data Visualization Studio


The McCasland Foundation Data Visualization Studio is a new addition to the Edmon Low Library. This room, also called 109A, is located on the western wall of first floor near Cafe Libro. The data visualization studio is open to anyone and is encouraged for projects. The studio holds a wall-mounted SMART Board as well as a tabletop monitor that allows collaboration between two screens. You can send information from one board to the other or reflect one screen to its counterpart.


The tabletop screen is great for close-up visuals of graphics and data. Data visualization itself is the display of information for analysis and communication, and the technology in the library’s studio is meant to do just that. The horizontal screen allows users to get closer to their data and see more detail. It is also useful to be able to move elements and create different arrangements of data.


Although it is set up horizontally, the data visualization screen is not an actual table, so food and beverage are not meant to be placed on it. You can reserve the studio through the Edmon Low Creative Studios.


Event Room


The new event room, also known as 102G, is primarily dedicated to hosting small events in the Edmon Low Library. It is located in the center of the new 1st floor study rooms, and the space can fit up to 40 people, depending on seating arrangement. With a SMART Board, wireless controls and a podium, this room is ideal for hosting seminars or holding presentations.  


When it is not being used for an event, 102G is open to the public for studying. There have been instances of visitors being asked to leave the room for an event,but this conflict can be avoided by checking the event calendar. The room can be reserved by completing a request form online.


Reading Room


The Anne Morris Greenwood Reading Room is one of the library’s oldest study areas. The space has been used by students since the 1950's, and it still attracts the most diligent of scholars today. It is the only space in the library to offer lamps, which makes it a spot of interest for students who perform their best with personal space and extra amenities.


It has two levels, which is convenient for those who want to be productive during their library visit but do not wish to venture higher than the second floor. There are also more electric outlets in this room than any single room in the library.


Contrary to popular belief, group study is allowed and encouraged in the Reading Room; however, the room remains quiet much of the time. Events on the second floor are typically held in the Peggy V. Helmerich Browsing Room, so the Reading Room remains a dedicated study space around the clock.

You may still have to see these rooms for yourself to get the full effect of their worth, but there is no doubt these spaces are bringing intelligence and innovation to the OSU Library. Whether data development or diligent studying is the goal, the library remains a primary location for traditional services and advancing technology.

Post by Jennifer Barnett, E-communications intern for OSU Library Communications