From e02f14314b561c19ecaa61f7eb048b0f8588981f Mon Sep 17 00:00:00 2001 From: Nesar976 Date: Fri, 23 Jan 2026 00:03:05 +0530 Subject: [PATCH 1/3] Add redirects for marine life data network documentation pages --- _docs/mbon-data-flow.md | 1 + _docs/metadata-eml.md | 1 + 2 files changed, 2 insertions(+) diff --git a/_docs/mbon-data-flow.md b/_docs/mbon-data-flow.md index 6def081..c26a613 100644 --- a/_docs/mbon-data-flow.md +++ b/_docs/mbon-data-flow.md @@ -1,5 +1,6 @@ --- title: "Data Flow" +redirect_to: https://ioos.github.io/marine_life_data_network/data-flow.html keywords: data tags: [biology, dataflow] toc: true diff --git a/_docs/metadata-eml.md b/_docs/metadata-eml.md index ceae874..050100e 100644 --- a/_docs/metadata-eml.md +++ b/_docs/metadata-eml.md @@ -1,5 +1,6 @@ --- title: "How-to guide for MBON metadata" +redirect_to: https://ioos.github.io/marine_life_data_network/metadata-eml.html keywords: metadata toc: true tags: [metadata, eml] From 77f9a6ef9fa286b093319efd3ea3c94811d93220 Mon Sep 17 00:00:00 2001 From: Nesar976 Date: Tue, 27 Jan 2026 12:00:00 +0530 Subject: [PATCH 2/3] docs: replace redirect_to with plugin-free meta redirects --- _docs/mbon-data-flow.md | 5 +++-- _docs/metadata-eml.md | 4 +++- 2 files changed, 6 insertions(+), 3 deletions(-) diff --git a/_docs/mbon-data-flow.md b/_docs/mbon-data-flow.md index c26a613..a5ceb4a 100644 --- a/_docs/mbon-data-flow.md +++ b/_docs/mbon-data-flow.md @@ -1,12 +1,13 @@ --- -title: "Data Flow" -redirect_to: https://ioos.github.io/marine_life_data_network/data-flow.html + keywords: data tags: [biology, dataflow] toc: true summary: This is a summary of the Marine Biodiversity Observation Network (MBON) data flow for Tabular data and metadata. mermaid: true --- + + {% include warning.html content="As of 2024-09-13 the content on this website is currently being migrated to the Marine Life Data Network website , please refer to the content there." %} diff --git a/_docs/metadata-eml.md b/_docs/metadata-eml.md index 050100e..2497207 100644 --- a/_docs/metadata-eml.md +++ b/_docs/metadata-eml.md @@ -1,11 +1,13 @@ --- title: "How-to guide for MBON metadata" -redirect_to: https://ioos.github.io/marine_life_data_network/metadata-eml.html + keywords: metadata toc: true tags: [metadata, eml] summary: This is a how-to guide for collecting MBON metadata. --- + + {% include warning.html content="As of 2024-09-13 the content on this website is currently being migrated to the Marine Life Data Network website , please refer to the content there." %} From fe59160b5b088111b1b24009e93b6c2d8b58f64b Mon Sep 17 00:00:00 2001 From: Nesar976 Date: Wed, 28 Jan 2026 23:43:49 +0530 Subject: [PATCH 3/3] docs: remove legacy MBON pages after migration --- _docs/data.md | 145 ----------------------- _docs/index.md | 46 -------- _docs/mbon-data-flow.md | 254 ---------------------------------------- _docs/metadata-eml.md | 144 ----------------------- _docs/metadata.md | 74 ------------ 5 files changed, 663 deletions(-) delete mode 100644 _docs/data.md delete mode 100644 _docs/index.md delete mode 100644 _docs/mbon-data-flow.md delete mode 100644 _docs/metadata-eml.md delete mode 100644 _docs/metadata.md diff --git a/_docs/data.md b/_docs/data.md deleted file mode 100644 index 588db54..0000000 --- a/_docs/data.md +++ /dev/null @@ -1,145 +0,0 @@ ---- -title: "Data Formatting" -keywords: data -tags: [biology, data] -toc: true -summary: This is Marine Biodiversity Observation Network (MBON) data recommendations. ---- - -{% include warning.html content="As of 2024-09-13 the content on this website is currently being migrated to the Marine Life Data Network website , please refer to the content there." %} - -# Data and File Formatting - -In choosing a file format, data collectors should select a format that is usable, open, and that -will likely be readable well into the future. Microsoft Excel, as an example, is a useful tool for -data manipulations and data visualization, but versions of Excel files may become obsolete and may -not be easily readable over the longer term. Likewise, database management systems (DBMS) like MS -Access, Filemaker Pro, and others, can be a very effective way to store and query data, but the raw -formats tend to change over time (even a few years). If your program or organization has used these -or other proprietary DBMS tools, it is essential to plan for exporting your data in a stable, -well-documented, and non-proprietary format. - -Below is a summary of the suggested tabular, image, and GIS data file formats suitable for long-term -archiving. - -* Containers: TAR, GZIP, ZIP -* Databases: CSV, XML -* Tabular data: CSV -* Geospatial vector data: SHP, GeoJSON, KML, DBF, NetCDF -* Geospatial raster data: GeoTIFF/TIFF, NetCDF, HDF-EOS -* Moving images: MOV, MPEG, AVI, MXF -* Sounds: WAVE, AIFF, MP3, MXF -* Still images: TIFF, JPEG 2000, PDF, PNG, GIF, BMP -* Text: XML, PDF/A, ASCII, UTF-8 - -For more complete guidance on best practices and appropriate formats for long-term preservation and -future accessibility, see the -[Library of Congress’ Sustainability of Digital Formats](https://www.loc.gov/preservation/digital/formats/) -web site, the LOC’s page on [Recommended Format Specifications for -preservation](https://www.loc.gov/preservation/resources/rfs/), and Hook et al’s recommendations in -[Best Practices for Preparing Environmental Data Sets to Share and -Archive](https://daac.ornl.gov/PI/BestPractices-2010.pdf). - -## Tabular Data - -Tabular data, or data in tables or spreadsheets, is by far the most common format for presenting, -analyzing, and storing data. However, common spreadsheet file formats are not ideal for sharing, -preserving, and reusing data; they’re not easily machine readable and may be difficult or impossible -to read if the specific software tools used to create them are significantly upgraded or outmoded. - -Using delimited text formats is a better way to ensure that tabular data are readable in the future. -A delimited text file is an ASCII-encoded file used to store data in which each line is uniquely -represented and had fields separated by a special character—the delimiter. Common delimiters are the -comma, tab, and colon. In a file with comma-separated values (a CSV file), the data values are -separated using commas as a delimiter. One benefit of being incredibly common and simple to use is -that most database and spreadsheet programs are able to read in or export data in a text-delimited -format. - -*Comma Separated Values (CSV)* is the most common delimited data format. If your data has strings or -sentences that may contain commas, be sure to wrap all string values in quotation marks (“”) or -another delimiter that may not be present in your data. The semicolon, (;), and a tab (“ “ , or “t”) -are other common delimiters. Text delimited files can be imported and exported by almost any -software designed for storing or manipulating data, including relational database systems, -spreadsheet software, and statistical analysis software. - -*ASCII (American Standard Code for Information Interchange)* is the most common text encoding and -the one most likely to be readable by tools. Other text encodings, such as UTF-8 are possible and -may be necessary for some non-English applications. Avoid obscure text encodings. Use ASCII if -possible, with UTF-8 or UTF-16 as secondary options. - -*Relational database management systems (RDBMS)* (such as Microsoft Access) create file formats that -need specialized software to open and view the contained information. Creating CSV files or ASCII -text versions and PDF/A’s of the data provider’s original data ensures that the information -contained within the file is openly accessible to data customers. Tabular data stored within -relational databases should be broken out into CSV files or ASCII text versions with all table -relationships described in enough detail for the database to be recreated. - -## Data Headers - -In order for others to use your data, they must fully understand the contents of the dataset. Most -commonly, data values within a file are organized in rows and columns, with each event or -observation as a row and each column representing a measurement or contextual information. Use a -header row at the top of each file describing the values in column. - -Follow these best practices for data headers: -* Use commonly accepted parameter names for header titles (e.g. site, date, treatment, -units_of_measure, etc). -* Use consistent capitalization of header names (e.g. temp and precip, or Temp and Precip). -* Explicitly state units of reported parameters in the data file and the metadata. -* If a coded value of abbreviation is used, be sure to provide a definition in the metadata. -* Adopt a similar structure across data files. If the same parameters are used across multiple -files, use a file template to maintain consistent column names across files. Avoid having different -numbers of columns or rearranging columns across similar files. -* Column names or headings should contain only numbers, letters, hyphens, and underscores—no spaces -or special characters. This also applies to column names and table names in databases. Special -characters and spaces are error-prone when machine-read or edited. - -## Data Values - -To allow others to best understand your data, follow these best practices for values within a data -file: -* Standardize all coded and null values within a dataset. -* Use an explicit value for missing or no data, rather than an empty field. Or, distinguish between -a zero and a blank value for numeric fields. -* For numeric fields, represent missing data with a specified extreme value (e.g., -9999), the IEEE -floating point NaN value (Not a Number), or the database NULL. Be advised that NULL and NaN can -cause problems, particularly with some older programs. For character fields, use NULL, “not -applicable”, “n/a” or “none”. -* If there are multiple reasons that cells might not have values, include a separate code for each -reason. -* The null value(s) should be consistently applied within and among data files. -* If data values are encoded, be sure to provide a definition in the metadata. -* Don’t include rows with summary statistics. It is best to put summary statistics, figures, -analyses, and other summary content into a separate companion data file. - -## Specific Formatting Recommendations for Biological Data - -Biological data is often stored in tabular formats such as spreadsheets and database tables. -Examples of biological datasets may include environmental data, such as temperature, salinity, or -conductivity, but focus on measuring variables related to one or more species of plant or animal. -Examples of biological data include, but are not limited to, marine bird surveys, genetic analyses -of salmon stocks, and toxicological analyses of lichen. - -For biological data, the following best practices apply: - -1. Align to [Darwin Core](https://dwc.tdwg.org/), when possible. -1. Archive data in CSV (or another non-proprietary, text-based format) whenever possible. -2. Follow established conventions when naming variables or columns. Refer to the [List of Darwin Core -Terms](https://dwc.tdwg.org/list/) -3. Define distinct events (such as location, time (with time zone), and/or depth) within a file with -a unique identifier. The identifier is often presented as sample_id or collection_event_id. See -Darwin Core term [eventID](https://dwc.tdwg.org/list/#dwc_eventID) -4. Include both the common name, the scientific name, and the -[WoRMS AphiaID code](http://www.marinespecies.org/aphia.php?p=match) for each species. The -[ITIS Taxonomic Serial Number (TSN)](https://www.itis.gov/) is another option if WoRMS does not meet -the needs of your data. - -## Additional Data Documentation - -* Save documentation about the data in non-proprietary file formats, such as .txt, .xml, or .pdf. -* Images, pictures, or figures should be saved as JPEG or GIF files. -* The name of the documentation should follow a logical naming convention identical to the related -data file(s), but indicating that the file is a metadata record, e.g. `[data_file_name]_METADATA.xml`. -* For complicated datasets, supplemental documentation is more useful when structured as a user’s -guide for the data. When constructing such a guide, include enough detail for someone with -sufficient domain knowledge to understand, trust, and reuse your data 20+ years in the future. diff --git a/_docs/index.md b/_docs/index.md deleted file mode 100644 index 899e4e8..0000000 --- a/_docs/index.md +++ /dev/null @@ -1,46 +0,0 @@ ---- -title: "MBON Data and File Formatting" -keywords: homepage -tags: [getting_started, about, overview] -toc: false -#permalink: index.html -summary: This documentation describes the Marine Biodiversity Observation Network (MBON) data and file formatting recommendations. ---- - -{% include warning.html content="As of 2024-09-13 the content on this website is currently being migrated to the Marine Life Data Network website , please refer to the content there." %} - -## Introduction - -[Marine Biodiversity Observation Network (MBON)](https://marinebon.org) observation data is focused on organisms from microbes to whales, including measures of biodiversity (e.g. presence, abundance), productivity, genomics, phenology, and other relevant ecological process measurements or indices. Also featured are habitat characterization and habitat diversity measures, including satellite data and added-value data derived from satellite observations, and neural network model results, such as biogeographical seascape classifications. - -The data have been generated within the MBON regions of the Arctic, Central California, Southern California, the Gulf of Maine, the Pacific Northwest, and South Florida. Data have been collected by associated scientists or provided by multiple other independent programs, such as the IOOS Regional Associations, Long-Term Ecological Research (LTER) programs, universities, and other fisheries or marine wildlife institutions. - -This website describes the recommendations for formatting and sharing data and metadata for the MBON community. The materials presented here were developed through the MBON Data Management and Cyberinfrastructure Working Group (MBON DMAC WG). The working group charter can be found [here]({{ site.url }}/mbon-docs/working-group-charter.html). If you would like to contribute to this documentation, see [CONTRIBUTING.md](https://github.com/ioos/mbon-docs/blob/gh-pages/CONTRIBUTING.md). - -## Categories of MBON observations -- Taxonomic data - - See the [MBON data flow](https://ioos.github.io/mbon-docs/mbon-data-flow.html) for sharing and standardizing data to Darwin Core. -- Organisms abundance - - See the [MBON data flow](https://ioos.github.io/mbon-docs/mbon-data-flow.html) for sharing and standardizing data to Darwin Core. -- Genetic make-up (‘omics, including informatics requirements) - - For eDNA, please see the [NOAA Omics Data Management Guide](https://noaa-omics-dmg.readthedocs.io/en/latest/) as the authoritative source for proper data management for MBON projects and the IOOS community. -- Acoustics (active and passive) - - For passive acoustic monitoring, please see [NCEI's Passive Acoustic Data Best Practices](https://www.ncei.noaa.gov/products/passive-acoustic-data#tab-3561) as the authoritative source for proper data management for MBON projects and the IOOS community. -- Imaging -- Optics -- Animal tracking - - For satellite telemetry data, please see the [Integrated Ocean Observing System (IOOS) Animal Telemetry Network Data Assembly Center (ATN DAC)](https://atn.ioos.us/help/) as the authoritative source for proper data management for MBON projects and the IOOS community. - - - -## Website contents -- [MBON Data Flow]({{ site.url }}/mbon-docs/mbon-data-flow.html) - This is a summary of the Marine Biodiversity Observation Network (MBON) data flow. -- [Data and File Formatting]({{ site.url }}/mbon-docs/data.html) - This is Marine Biodiversity Observation Network (MBON) data recommendations. -- [Metadata and Documentation]({{ site.url }}/mbon-docs/metadata.html) - This is Marine Biodiversity Observation Network (MBON) metadata recommendations. -- [How to guide for MBON Metadata]({{ site.url }}/mbon-docs/metadata-eml.html) - This is a how-to guide for collecting MBON metadata. -- [MBON use case]({{ site.url }}/mbon-docs/use-case.html) - This is a collection of IOOS Marine Biodiversity Observation Network (MBON) data flow use cases. -- [MBON Data Portal Demo]({{ site.url }}/mbon-docs/data-portal-demo.html) - This page demonstrates the MBON Data Portal and some of the tools in the portal. - ---- - -For help with the [MBON Data Portal](https://mbon.ioos.us), please see the [MBON portal help documentation](https://mbon.ioos.us/help/). diff --git a/_docs/mbon-data-flow.md b/_docs/mbon-data-flow.md deleted file mode 100644 index a5ceb4a..0000000 --- a/_docs/mbon-data-flow.md +++ /dev/null @@ -1,254 +0,0 @@ ---- - -keywords: data -tags: [biology, dataflow] -toc: true -summary: This is a summary of the Marine Biodiversity Observation Network (MBON) data flow for Tabular data and metadata. -mermaid: true ---- - - - -{% include warning.html content="As of 2024-09-13 the content on this website is currently being migrated to the Marine Life Data Network website , please refer to the content there." %} - -# MBON Data Flow - -```mermaid -%%{ - init: { - 'theme': 'base', - 'themeVariables': { - 'primaryColor': '#007396', - 'primaryTextColor': '#fff', - 'primaryBorderColor': '#003087', - 'lineColor': '#003087', - 'secondaryColor': '#007396', - 'tertiaryColor': '#CCD1D1' - }, - 'flowchart': { 'curve': 'basis' } - } -}%% - -flowchart TD - -A["Tabular Data -& -Metadata"] - -B[("Raw Data -Access Point -(eg. RA ERDDAP)")] - -C("Darwin Core -Alignment") - -D[(NCEI)] - -E[("IPT -OBIS-USA")] - -F[/"MBON -Data Portal"\] - -G([OBIS]) - -H([GBIF]) - -I[("IOOS Data Catalog -(data.ioos.us)")] - -J[(NOAA OneStop)] - -K[(data.gov)] - -L[("Commerce -Data Hub")] - -M[/"IOC-UNESCO Harmful Algae Information System"\] - -N[/"Infographics"\] - -A --> B -B --> C -E --> D -B --> D -C --> E -E --> G -E --> H -B -- If using IOOS RA ERDDAP --> I - -D --> FC -I --> FC - -H .-> EP -G .-> EP -B .-> EP -D .-> EP - - -subgraph EP [Example Products] -M -N -F -end - -subgraph FC [Federal Catalogs] -J -K -L -end - -click D "https://www.ncei.noaa.gov" "NCEI" _blank -click F "https://mbon.ioos.us" "MBON" _blank -click G "https://obis.org" "OBIS" _blank -click H "https://gbif.org" "GBIF" _blank -click I "https://data.ioos.us" "IOOS Catalog" _blank -click J "https://data.noaa.gov/onestop/" "NOAA OneStop" _blank -click K "https://data.gov" "data.gov" _blank -``` - - -For data collected/managed by an IOOS MBON project, the project should ensure data and information are readily available -to resource managers, scientists, educators, and the public in an easily digestible way. To that end, coordination with -an IOOS Regional Association to make the data available via ERDDAP services meets these goals. Using the services that -ERDDAP provides, a data manager can develop a reproducible workflow for aligning the data to the -[Darwin Core standard](https://dwc.tdwg.org/). Finally, submission to NCEI ensures that no observations are lost and -there is long-term stewardship of these data, as well as meeting our PARR requirements. The sections below provide more -context as well as tips and tricks for each of the elements in the diagram above. - -## RA ERDDAP -For the IOOS MBON projects ERDDAP is used as a mechanism for quickly and efficiently sharing biological observations with -the broader community. While ERDDAP can provide data access following the FAIR principles, further alignment to Darwin -Core and submission to OBIS is necessary to make these observations more useful to a broader audience. Essentially, serving -data through an RA ERDDAP is one part of a larger process and should be treated as such. - -### Key principles for data -When preparing a dataset to be served via ERDDAP it is recommended to follow a few key principles for data management. -* For organizing your data files, follow the [Tidy data](https://r4ds.had.co.nz/tidy-data.html) recommendations: - * Variables as columns - * Observations as rows - * Don't embed data in the column headers. -* Follow [ISO-8601](https://en.wikipedia.org/wiki/ISO_8601) for dates - * `YYYY-MM-DDTHH:mm:ssZ` (eg. 2021-08-19T12:38:22Z) - * Include the time zone. -* Latitude and Longitude in decimal degrees (WGS84 preferred) -* Identify units of measure -* Check species names against [WoRMS](https://www.marinespecies.org/). -* See the sections on [Data and File Formatting](data.html) and -[Metadata and Documentation](metadata.html) for more recommendations and best practices. - -**Additional Resources** -* [Configuring datasets.xml](https://coastwatch.pfeg.noaa.gov/erddap/download/setupDatasetsXml.html) - The ERDDAP manual -for configuring a dataset. -* [ERDDAP Quick Start Guide](https://ioos.github.io/erddap-gold-standard/) - Quick Start Guide for deploying ERDDAP in a Docker Container. -* [ERDDAP Google Group](https://groups.google.com/g/erddap) - A great place to search for questions and ask your -questions. - -### ERDDAP Requirements - -Below is a list of the absolute bare minimum pieces of metadata required by ERDDAP. Some dataset types might have other -requirements specific to the data file formats. -* Global attributes - * [datasetID](https://coastwatch.pfeg.noaa.gov/erddap/download/setupDatasetsXml.html#datasetID) - * [sourceUrl](https://coastwatch.pfeg.noaa.gov/erddap/download/setupDatasetsXml.html#sourceUrl) - however, depends on -dataset type - * [infoUrl](https://coastwatch.pfeg.noaa.gov/erddap/download/setupDatasetsXml.html#infoUrl) - * [institution](https://coastwatch.pfeg.noaa.gov/erddap/download/setupDatasetsXml.html#institution) - * [summary](https://coastwatch.pfeg.noaa.gov/erddap/download/setupDatasetsXml.html#summary) - * [title](https://coastwatch.pfeg.noaa.gov/erddap/download/setupDatasetsXml.html#title) -* Variable attributes - * [dataVariable](https://coastwatch.pfeg.noaa.gov/erddap/download/setupDatasetsXml.html#dataVariable) - * [sourceName](https://coastwatch.pfeg.noaa.gov/erddap/download/setupDatasetsXml.html#sourceName) - -### ERDDAP Tips and Tricks -* Date/time - It is not a requirement to have a variable assigned to `time` (ie. -`time`) in ERDDAP. If a variable's destinationName is set to `time`, ERDDAP will use -the `units` attribute to attempt to interpret the datum. See the documentation on [How ERDDAP Deals with -Time](https://coastwatch.pfeg.noaa.gov/erddap/convert/time.html#erddap) for more information. - * *Trick* - If you don't want the variable interpreted as a time, set the `` to something other than -`time`. For example, in your source file the coumn `time` has a value of `2020-01-01`, but you don't want that -interpreted by ERDDAP. Then, set the `destinationName` to `time2` and ERDDAP will treat the field as a string. - * *Caution* - If you do not have an assigned `time` variable in a dataset, some of the access formats might not be -available (eg. .esricsv, .odvtxt). -* Latitude/Longitude - Similar to date/time above, it is not a requirement to have latitude/longitude variables. However, -the dataset will have a limited amount of access formats. -* *Trick* - ERDDAP now has the capability to create derived variables from existing fields (since v2.10). See the -documentation on [Script SourceNames / Derrived -Variables](https://coastwatch.pfeg.noaa.gov/erddap/download/setupDatasetsXml.html#scriptSourceNames). -* *Trick* - ERDDAP can handle media files, such as image, audio and video files. See -[MediaFiles](https://coastwatch.pfeg.noaa.gov/erddap/download/setupDatasetsXml.html#MediaFiles) for more information. - * *Bonus Trick* - -[EDDTableFromFileNames](https://coastwatch.pfeg.noaa.gov/erddap/download/setupDatasetsXml.html#MediaFiles) allows you to -create a dataset from information about files in the file system. While it doesn't serve data from within the files it -does provide a mechanism for sharing data in other formats (eg. zip packages, Word docs, Excel spreadsheets, etc.). The -resultant dataset in ERDDAP is composed of the following columns: `url`, `name`, `lastModified`, and `size`. - -## Darwin Core alignment -When aligning a dataset to Darwin Core it is recommended that a data manager starts with serving the data via ERDDAP -or some comparable online system which has an [API](https://en.wikipedia.org/wiki/API) (or a way to programmatically -grab the data). When working through the Darwin Core alignment using a scripting language (eg. R or Python) which uses -the data served via ERDDAP (or comparable service) is highly recommended. A scripting language provides provenance, -transparency, and reproducibility for the translation. This helps reduce the amount of errors and back-and-forth between -data managers and OBIS. It is highly recommended that, if using a scripting language, the scripts are shared via -distributed version control systems like [GitHub](https://www.github.com). - -**Recommendations TL;DR;** -* Follow the guidance at [TDWG's Darwin Core quire reference guide](https://dwc.tdwg.org/terms/). -* Use a scripting language. -* Script should point to source data on a hosted web service. -* Scripts should be shared via GitHub. - -**Additional Resources** -* [Standardizing Marine Biological Data Guide](https://ioos.github.io/bio_data_guide/) - A guide and examples of -aligning datasets to Darwin Core. -* [Aligning data to Darwin Core notebook in IOOS CodeLab](https://ioos.github.io/ioos_code_lab/content/code_gallery/data_management_notebooks/2020-12-08-DataToDwC.html) - A Python notebook for aligning a dataset to Darwin Core available in the IOOS Code Lab. -* [OBIS Manual](https://manual.obis.org/) - This manual provides an overview on how to contribute data to OBIS and how to acess data from OBIS - -## Sending to OBIS-USA -Below are the various options for sending your data to OBIS-USA. - -* Attend the monthly [Standardizing Marine Biological Data Working Group](https://github.com/ioos/bio_data_guide#monthly-meetings) meeting and discuss transfer options. -* Contribute your dataset (and code) to the `datasets/` directory in the ioos/bio_data_guide repository ([here](https://github.com/ioos/bio_data_guide/tree/main/datasets)). See the [Contribute example applications](https://github.com/ioos/bio_data_guide/blob/main/CONTRIBUTING.md#contribute-example-applications) documentation for more information. -* Email Darwin Core aligned files to OBIS-USA: . -* Or, use the [Dataset Review Request](https://github.com/ioos/bio_data_guide/issues/new/choose) issue to initialize the request. - -## Sending to NCEI -When planning on submitting data to NCEI, the data provider should coordinate submissions through the IOOS Office to -identify which submission system should be used. This will ensure that the dataset is appropriately identified, tracked, -and stewarded through the submission process. - -Ideally, the raw data should be archived at NCEI. Typically, this will be the dataset served through -the [IOOS RA ERDDAP](#ra-erddap) and following the [key principles](#key-principles-for-data) laid out above. Archiving the dataset in its more raw -form (vs the Darwin Core aligned form) ensures that no information is lost. This also ensures that data providers -can always go back to the source data if issues arise. - -For more information about archiving data at NCEI, see [https://www.ncei.noaa.gov/archive](https://www.ncei.noaa.gov/archive). - -Briefly, the submission package sent to NCEI should indicate that the observations are from an IOOS MBON project (or -has some affiliation with IOOS). Below is a short summary of the two submission systems at NCEI and their intended uses. -* [ATRAC](https://www.ncdc.noaa.gov/atrac/guidelines.html) - Use the Advanced Tracking and Resource Tool for Archive -Collections (ATRAC) to submit repeating or multiple delivery data, or data that exceeds 20 GB. -* [S2N](https://www.nodc.noaa.gov/s2n/) - Use Send2NCEI to submit non-repeating or single delivery data less than 20 GB. - -**Note:** NCEI and OBIS-USA have established an automated process to archive the datasets from the OBIS-USA IPT. The process archives -the Darwin Core Archive version of the dataset and updates the NCEI Archival Information Package found at . While the OBIS-IPT is an extremely valuable product, the raw data should be -archived at NCEI as well. - -## Loading into MBON Portal -As depicted in the [data flow diagram](#mbon-data-flow), the MBON data portal can retrieve data from a variety of -sources. The two preferred sources for data include OBIS (or GBIF) and/or ERDDAP (hosted by a Regional Association), -however other web services could be acceptable to bring data in. In some cases, the MBON data portal might bring in -occurrence data through OBIS as well as additional observations that are served through ERDDAP. Below are the recommended -steps to load data into the MBON Portal: - -1. The dataset should be registered in the [MBON dataset registration form](https://docs.google.com/forms/d/e/1FAIpQLSfguACbLmcLiFxHKsR5W5Mv9nEfd0E8oX2rY78gdwAYTrq_zA/viewform?usp=sf_link). - 1. This will ensure that we are aware of the dataset and have identified next actions to take. - 2. Identify that you would like the dataset visualized in the MBON portal and include a description of what that -visualization might be. -2. Share the dataset through OBIS/GBIF and/or through ERDDAP. - 1. For OBIS/GBIF see [Darwin Core alignment](#darwin-core-alignment) and [Sending to OBIS-USA](#sending-to-obis-usa). - 2. For sharing through ERDDAP see [RA ERDDAP](#ra-erddap). -3. Iterate with the MBON portal development team to ensure the visualizations are appropriate for the observations. - -_To note_ There are additional pathways to share data with the MBON portal using the Research Workspace. For more -information on that pathway see [Contribute Data in the MBON portal documentation](https://mbon.ioos.us/help/how-to/catalog/contribute-data.html). diff --git a/_docs/metadata-eml.md b/_docs/metadata-eml.md deleted file mode 100644 index 2497207..0000000 --- a/_docs/metadata-eml.md +++ /dev/null @@ -1,144 +0,0 @@ ---- -title: "How-to guide for MBON metadata" - -keywords: metadata -toc: true -tags: [metadata, eml] -summary: This is a how-to guide for collecting MBON metadata. ---- - - - -{% include warning.html content="As of 2024-09-13 the content on this website is currently being migrated to the Marine Life Data Network website , please refer to the content there." %} - -The information below can also be viewed as a Microsoft Word document [here](https://github.com/ioos/mbon-docs/raw/gh-pages/assets/EML.Metadata.Template.docx). - -# EML Metadata -Highlighted elements are required -
-
-**Dataset shortname (will be the URL in the IPT, all lower case, no spaces):**
-**Title:**
-**Update Frequency (pick one):** Daily, Weekly, Monthly, Biannually, Annually, As Needed, Continually, Irregular, Not Planned, Unknown, Other Maintenance Period
-**Data License (pick one):** CC-0, CC-BY, CC-BY-NC
-**Description / Abstract:**
-**Resource Contacts (include First Name, Last Name, Position, Organization, Email, can link to ORCID):**
-**Resource Creators (include First Name, Last Name, Position, Organization, Email, can link to ORCID):**
-**Metadata Providers (include First Name, Last Name, Position, Organization, Email, can link to ORCID):**
- -## Geographic Coverage -**Set global coverage: YES or NO (if NO rest is required)**
-**Description:**
-**West (decimal degrees):**
-**East (decimal degrees):**
-**South (decimal degrees):**
-**North (decimal degrees):**
- -## Taxonomic Coverage -**Description:**
-Scientific Name, Common Name, Taxon Rank for all species or at higher level if all part of one higher level rank - -## Temporal Coverage -**Temporal Coverage Type (can select more than one but corresponding information must align with the type):** Single Date, Formation Period, Date Range, Living Time Period
- -## Keywords -**Thesaurus / Vocabulary:**
-**Keyword List:**
- -## Associated Parties - -All `US MBON` affiliated datasets should include, at least, one **Associated Party** which is affiliated with `US MBON`. This allows OBIS to create an institute page which reflects `US MBON`'s contributions (similar to SCB MBON ). This is possible because `US MBON` has an OceanExpert institution which will link the various datasets together in an institute page. - - -US MBON Ocean Expert Institution:
-US MBON OBIS Institution: - -_To create an **OceanExpert** institution, follow the -[Appendix: How-to Create an OceanExpert Institution](#appendix-how-to-create-an-oceanexpert-institution)._ - -**Associated Party (include First Name, Last Name, Position, Organization, Email, can link to ORCID):**
-**Associated Party Role (pick one per person from table below):**
-_`Role` information is also available at ._ - -| **Role** | **Description** -|:----------------------|:------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ -| author | an agent who is an author of a publication that used the dataset, or author of a data paper -| contentProvider | an agent who contributed content to a dataset (the dataset being described may be a composite) -| custodianSteward | an agent who is responsible for/takes care of the dataset -| distributor | an agent involved in the publishing/distribution chain of a dataset -| editor | an agent associated with editing a publication that used the dataset, or a data paper -| metadataProvider | an agent responsible for providing the metadata -| originator | an agent who originally gathered/prepared the dataset -| owner | an agent who owns the dataset (may or may not be the custodian) -| pointOfContact | an agent to contact for further information about the dataset -| principalInvestigator | a primary scientific contact associated with the dataset -| processor | an agent responsible for any post-collection processing of the dataset -| publisher | the agent associated with the publishing of some entity (paper, article, book, etc) based on the dataset, or of a data paper -| user | an agent that makes use of the dataset -| programmer | an agent providing informatics/programming support related to the dataset -| curator | an agent that maintains and documents the specimens in a collection. Some of their duties include preparing and labeling specimens so they are ready for identification, and protecting the specimens -| reviewer | person assigned to review the dataset and verify its data and/or metadata quality. This role is analogous to the role played by peer reviewers in the scholarly publication process. - -_Note: Abby Benson will be listed as the `publisher` if she loads the data into the IPT for you and publishes to GBIF and/or OBIS_ - -_Note: Mathew Biddle will be listed as the `distributor` with `US MBON` as the institution._ - -## Project Data -_If including this section Title and Project Personnel are required_ - -**Title:**
-**Identifier**
-**Description:**
-**Funding:** This work was supported by the U.S. Marine Biodiversity Observation Network (MBON) co-organized by NOAA, NASA, BOEM, and ONR through the National Oceanographic Partnership Program (NOPP) [(add AGENCY grant # here if required/preferred)]
-**Study Area Description:**
-**Design Description:**
-**Project Personnel (First Name, Last Name, Role):**
- -## Sampling Methods -_If including this section Study Extent, Sampling Description, Step Description are required)_ - -**Study Extent:**
-**Sampling Description:**
-**Quality Control:**
-**Step Description:**
- -## Citations -**Resource Citation (can be autogenerated):**
-**Resource Citation Identifier:**
-**Bibliographic Citations:**
- -## Collection Data -**Collections:**
-**Specimen Preservation Methods:**
-**Curatorial Units:**
- -## External Links -**Resource Homepage:**
-**Other Data Formats (Name, Character Set, Download URL, Data Format, Data Format Version):**
- -## Additional Metadata -**Resource Logo URL (or file upload):**
-**Purpose:**
-**Maintenance Description:**
-**Additional Information:**
-**Alternative Identifiers:**
- -## Appendix: How to create an OceanExpert institution - -Before creating a new institution it's always good to first do a search to make sure your institution doesn't already -exist in **OceanExpert**. You can search for institutions at . Be sure to check -various spellings and acronyms too! - -To create an **OceanExpert** institution: -1. Create an **OceanExpert** account: -2. Once your account is created, log in and edit your profile:
-![image](https://user-images.githubusercontent.com/8480023/207622140-89c3bbb5-1abd-4058-809b-8434635bfa54.png) -3. In the **Affiliation and Address** section of the profile there is a dropdown menu for `Organisation/Institution/Company`, select the first option `Cannot find my Institute - Create New`:
-![image](https://user-images.githubusercontent.com/8480023/207622420-52592ce2-bda7-4ba4-a40a-47417ca6bca7.png) -4. A new window will pop up and ask you to populate a new institution record:
-![image](https://user-images.githubusercontent.com/8480023/207622801-e395e6af-dd61-4247-a865-b5a86dcaf787.png) - 1. The only required fields are: `Name`, `Type` (choose from Academic, Government, International / Intergovernmental, NGO, Private Commercial, Private non-profit, and Research), `Address`, `Zip Code`, and `City`. -5. Click **Add Institution**. -6. That should add the institution to the drop down list for `Organisation/Institution/Company` so that you can associate it with your profile. - -To edit **OceanExpert** institution pages, send an email to with the requested changes. diff --git a/_docs/metadata.md b/_docs/metadata.md deleted file mode 100644 index 430c521..0000000 --- a/_docs/metadata.md +++ /dev/null @@ -1,74 +0,0 @@ ---- -title: "Metadata Formatting" -keywords: metadata -tags: [biology, metadata] -toc: true -summary: This is Marine Biodiversity Observation Network (MBON) metadata recommendations. ---- - -{% include warning.html content="As of 2024-09-13 the content on this website is currently being migrated to the Marine Life Data Network website , please refer to the content there." %} - -# Metadata and Documentation - -Descriptive metadata and documentation are critical to maintaining data quality. Metadata is “data -about the data” that describes and contextualizes the dataset to ensure it is understandable to -future users. Beyond standardized metadata, useful documentation might include standard operating -procedures, field notes, etc., from which metadata may be derived or referenced. - -Throughout the data lifecycle, both the metadata and documentation must be recorded and updated to -reflect the actions taken to the data. This includes collection, acquisition, processing, quality -review, and analysis, as well as any other stage of the data lifecycle. - -The dataset’s metadata and/or documentations (or a link to it) must be distributed with the dataset. - -## Metadata - -Metadata describes information about a dataset to ensure that it can be understood and re-used -properly in the future. Content of the metadata record includes where the data were collected, who -is responsible for the data, why the data were created, how the data are organized, and any derived -data calculation methods (e.g. for fish abundance estimate) that were used. Metadata generally -follows a standard format to ensure the semantics of metadata fields are understood by creators and -consumers of the metadata, to ease the use of the metadata in catalog and discovery systems, and -simplify automatic machine-to-machine transfer of records. - -There are two recomended metadata formats. -1. Ecological Markup Language (EML) - [https://eml.ecoinformatics.org/](https://eml.ecoinformatics.org/) -2. ISO 19115 family of standards - [https://www.fgdc.gov/metadata/iso-standards](https://www.fgdc.gov/metadata/iso-standards) - -### EML -The EML project is an open source, community oriented project dedicated to providing a high-quality metadata -specification for describing data relevant to diverse disciplines that involve observational research like ecology, -earth, and environmental science. The specification is maintained by voluntary project members who donate their time -and experience in order to advance information management for ecology. Project decisions are made by consensus of the -current maintainers on the project. - -See this documentation for more information: -Matthew B. Jones, Margaret O’Brien, Bryce Mecum, Carl Boettiger, Mark Schildhauer, Mitchell Maier, Timothy Whiteaker, -Stevan Earl, Steven Chong. 2019. Ecological Metadata Language version 2.2.0. KNB Data Repository. doi:[10.5063/F11834T2](https://dx.doi.org/10.5063/F11834T2) - - -### ISO -It is recommended to use the ISO 19115-2 XML for metadata standards following the ISO 19115 guidance from -[NCEI](https://www.ncei.noaa.gov/sites/default/files/2021-03/AB-GUID-02823_R1_Guidance%20for%20The%20NCEI%20Collection%20Level%20Metadata%20Template%20v1.1._0.pdf). -See also [NCEI's guidance on metadata](https://www.ncei.noaa.gov/resources/metadata). - -Methods for generating metadata: -* Various software packages exist to create EML metadata. See the list of tools in the [standardizing marine biological -data guide](https://ioos.github.io/bio_data_guide/tools.html#tools). -* If data are being served by ERDDAP, then ERDDAP will automatically generate the ISO XML documents. -* If using the [Research Workspace](https://researchworkspace.com/intro/) to submit data, an -integrated metadata editor is included to generate metadata in the FGDC-endorsed ISO 19110 and -19115-2 standards for geospatial metadata. Refer to the -[Metadata Best Practices](https://www.axiomdatascience.com/best-practices/MetadataBestPractices.html#metadata-best-practices) -section for help creating scientific metadata using the Research Workspace metadata editor. This -document provides field-by-field guidance on how to write high-quality metadata. - -## Additional Data Documentation - -* Save documentation about the data in non-proprietary file formats, such as .txt, .xml, or .pdf. -* Images, pictures, or figures should be saved as JPEG or GIF files. -* The name of the documentation should follow a logical naming convention identical to the related -data file(s), but indicating that the file is a metadata record, e.g. `[data_file_name]_METADATA.xml`. -* For complicated datasets, supplemental documentation is more useful when structured as a user’s -guide for the data. When constructing such a guide, include enough detail for someone with -sufficient domain knowledge to understand, trust, and reuse your data 20+ years in the future.