Reading data#
The opening of the files that contains the data are on your own. But there’s some care to take.
In this guide we assume that the data from the average weather stations and the data from the fluxes stations are into two separate files. The whole programm use the pandas.DataFrame object. So you should use pandas to open the files.
Let’s start with the data_reader.py script and the header:
from datetime import datetime
import pandas as pd
import numpy as np
from hibou import utils, config
datetimefor time parsingpandasfor opening and storing datanumpyfor vectorize operations herehibouprovides some tools to compute sinusoidal time parameters and MetPy calculations
Once the header is done, we can start by defining our variables to train the artificial neural network (ANN). This step allows hibou to access the defined variables during the preprocessing of the data. These variables names are stored into the config module and more precisely into config.Inputs and config.Targets classes:
# redefining variables names for inputs
config.Inputs.TEMPERATURE = "Ts_Avg"
config.Inputs.PRESSURE = "Pa_Avg"
config.Inputs.WIND_SPEED = "wind_speed"
config.Inputs.RAIN = "P_mm_Avg" # we can even define new variables
config.Inputs.MIXING_RATIO = "mr_Avg"
config.Inputs.RELATIVE_HUMIDITY = "RelHum_Avg"
# redefining variables names for targets
config.Targets.H = "sensible_flux"
config.Targets.LE = "latent_flux"
Then, just before opening files, we create a small function that cast a string into a datetime object:
def parse_date(timestamp: str):
"""Cast a string formatted like YYYY-MM-DD HH:MM:SS into a datetime object."""
return datetime.strptime(timestamp, "%Y-%m-%d %H:%M:%S")
Once this step is done, we can create two functions (one per file):
def open_weather_data():
data = pd.read_csv("data/weather_data.csv")
# parse dates
data[config.Inputs.DATETIME] = np.vectorize(parse_date)(data["date"])
# converting pressure from kPa to hPa
data[config.Inputs.PRESSURE] = data[config.Inputs.PRESSURE] * 10
# reconstruct wind components
data[config.Inputs.U] = - data[config.Inputs.WIND_SPEED] * np.sin(data["wind_dir"] * np.pi / 180)
data[config.Inputs.V] = - data[config.Inputs.WIND_SPEED] * np.cos(data["wind_dir"] * np.pi / 180)
# sinusoidal time parameters
(
data[config.Inputs.ANNUAL_SIN], # even if the columns doesn't exists, they will be created automatically
data[config.Inputs.ANNUAL_COS],
data[config.Inputs.DAILY_SIN],
data[config.Inputs.DAILY_COS]
) = utils.get_sinusoidal_time_parameters(data[config.Inputs.DATETIME].values)
return data
The second function is very similar to the first one:
def open_flux_data():
data = pd.read_csv("data/flux_data.csv")
# parse dates
data[config.Targets.DATETIME] = np.vectorize(parse_date)(data["date"])
return data