Albin O. Kuhn Library & Gallery - Staff Wiki
ScholarWorks ETD Load Procedures
Files for Combining ETD and Zotero Loads
Notices from Proquest that files are available.
Filezilla.
Proquest FTP login info.
ETD Directory on hard drive with pdf and xml subdirectories.
7-Zip.
Adobe Acrobat Standard. Modify Acrobat settings: When you're in Acrobat, go to edit, then preferences. Click on "Documents" in the left-hand column. In the main part of the pop-up, under PDF/A view mode, use the drop-down to select "never."
Computer configured to open XML files with WordPad (Right click an XML file and select "Open with" and then "Chose Program." Select WordPad, then click "Always use the selected program to open this kind of file.").
Editix XML Editor.
XSL file for reformatting the XML files, ETDConversionForDspace.xsl (attached above).
Microsoft Excel with the Developer tab enabled (Left click on the windows symbol and select "Excel Options." On the popular tab, check "Show developer tab in the Ribbon." Go to the Trust Center Tab. Click "Trust Center Settings." Click "Enable all Macros.").
The latest version of Python installed on your computer with required PIP packages installed. Python and Pip Package Installation
Python ETD Conversion Program (attached above).
Python program that compares the file names in columns B, C, and D and the actual file names in the pdf directory
The modified PySAF program (the one on Github won’t work on our ETDs) attached above.
Python collection Mapping program (attached above). User Documentation for Mapping Script
Python date check program (attached above).
Python find empty collections program (attached above).
For converting video files to mp4's: Avidemux.
FTP and Unzip the files (about 100 at a time): Downloaded files through
After FTPing the files, be sure to delete them from the FTP server.
Use Filezilla to FTP the new thesis and dissertations from Proquest. Open Filezilla. Enter the Proquest FTP IP, the username, and port. Push Enter. The last line of the top box on the screen should say "Directory Listing Successful" and the lower left-hand portion of the screen should be populated with files on the Proquest server. The left side of the screen shows your computer--find the ETD folder on your hard drive. Use the date to identify the new files that we need to obtain. Highlight all of the files we need by holding down the shift key while clicking the first and last files you want highlighted. Drag them to your ETD folder. The progress of file transfer will show on the bottom of the screen. Wait while all files transfer (you can minimize and do something else). After downloading the files, be sure to move the downloaded files to trash.
Verify that you have all of the files that have been sent by checking the number of files against the number of files the e-mail notices said were successfully downloaded. Add the number of successful downloads. Highlight all the ETD files, right click, and select properties. The number of files stated in the notices should match the total here.
Use 7-Zip to extract the zip files. Open 7-Zip. The folder that your files are in should be selected in the bar across the top of the window. If not, use the drop-down arrow to find it. Once you are on the correct folder, all of your zip files should display in the window. Highlight all of the zip files by holding the the shift key by clicking the first and last files you want highlighted. Click extract. The destination for the extract opens to C:\ETD\ZIP*\. Delete the *\ so that the files all go into the ZIP folder. Click ok.
Use Windows Explorer to sort the files by going to the "View" menu and selecting "arrange by file type.". Select all of the files of a given type, and move the to the appropriate sub-folder: Highlight all of the PDF files by holding down the shift key while clicking the first and last files you want highlighted. Drag them to the PDF subfolder and drop them there (or alternately, copy and paste them). Highlight all of the XML files by holding down the shift key while clicking the first and last files you want highlighted. Drag them to the XML subfolder and drop them there or alternately, copy and paste them).
Transform the Metadata
Combine the XML files into 1 File
DOS prompt:
Click the Windows Start button and type .cmd in the box. Push enter. A box with DOS will open.
Change the directory to the where you want the new file to go by entering cd followed by the path for the directory. For example, “CD C:\ETD” changes the directory to the ETD directory. To go up one level, "CD .." To go to the root directory, "cd /"
To copy the individual xml metadata files, use copy path *.xml newfilename. For example, if your xml files are in the ETD\xml\ directory, “copy c:\ETD\xml\*.xml combined.xml”.
Notepad:
Open the new file in notepad. Copy <?xml version="1.0" encoding="iso-8859-1"?> from the beginning of the file. Find and replace with nothing by pasting it <?xml version="1.0" encoding="iso-8859-1"?> . Put the <?xml version="1.0" encoding="iso-8859-1"?> back at the beginning of the file, inserting a line break between it and the remainder of the XML.
Add <ETD> after the <?xml version="1.0" encoding="iso-8859-1”?> at the beginning with a line break between it and the remainder of the XML.
At the end of the file, add a line break and </ETD> at the end.
Save and close the file.
Reformat the XML File using Editix:
Open Editix.
Open the XSL file ETDConversionForDspace.xsl (go to file, open, then change the file type to XSLT 2.0 document (*.xsl *.xslt)
Go to XSLT/Xquery transform a document
In XML source find your XML file.
In result find the directory you want the new file to go in and type the name with the extension .xml
Click ok.
Prepare the metadata in Excel:
Create a new Excel for the set.
Go sheet1, cell 1A. Use developer import to import the file you created using Editix.
Save your spreadsheet and close it.
Find the python ETD program and double click it. Wait a few seconds for popup to appear.
Select your Excel file.
Click “Run Filter”.
Data Checks
Open the Excel file which should now have a 2nd sheet where the processed data appears. Go to it.
Sort by departments column. Check for any department that didn't fill in and check that they’re the correct department names using the collections file spreadshreet
Marine-Estuarine Environmental Sciences |
Prepare the files:
Check for anything unexpectedly left over. ETD's with extra files will unzip into folders or into additional zip files. These will require some initial manipulation to prepare them. Take a look at the extra files and handle each case as appropriate as follows:
Approval sheets--These are forms that the adviser signs approving the thesis or dissertation. These are extraneous extra files that should simply be deleted. Move the pdf and xml files to your usual directories and process as usual.
Other data that is usually included in the main file and pdf appendices--Combine the extra files with the main file using Adobe Acrobat.Move the files to your usual directories and process as usual.
Non-pdf appendices and other files not meeting above criteria--Convert them to the most appropriate file format given here: Non-Proprietary File Formats (if not already in one of these formats). If it can't be satisfactorily converted to one of those file formats, leave it in the format that it's in. Put these in a supplement folder in your pdf folder.
Open each PDF. IF there is personal information such as phone numbers or addresses included in the CV, delete it. Otherwise, leave it in the document. Some files will have missing thesis or dissertation. Handle these as follows:
Missing Documents There are two reasons a thesis or dissertation may be missing. The document may be embargoed, or the document may have not been FTP'ed because it's a large file that couldn't be sent via the Proquest administration page, so was sent to Proquest on disk. To determine which case this is, take a look at the metadata and the DISS_submission publishing_option tag. This is usually the first field in the metadata. In that tag, there is a an embargo code set with a numeric embargo code:
"0" - No embargo
"1" - 6 month embargo
"2" - 1 year embargo
"3" - 2 year embargo
"4" - Until specified date
If the code is 0, we should have the file, and can obtain it by downloading from the ContentDM Administrator Resources & Guidelines page at http://www.etdadmin.com/cgi-bin/main/resources?siteId=75. Click on Dissertations & Theses @ University of Maryland, Baltimore County and search for the missing document. When you find it, download it and process as usual.
If the code is 1-4, the document is embargoed and we won't receive the document until the embargo period has passed.
At the end of the metadata file there is a DISS_sales_restriction code," and the date in that tag indicates when the embargo will expire and when we should receive that file. Note the file name along with the date the embargo will expire in the embargo list at the end of this procedure so that we can ensure that we receive the file when the time comes. When you process the metadata for embargoed documents in Excel, insert a note into the metadata for the document stating: "At the author's request, this dissertation isn't being made available at this time." The metadata is then uploaded as usual along with the title page. The metadata will be revised to remove this note when we receive the full file.
For other problems with the files Proquest FTP's to us, ask Michelle to call Proquest technical support at 877-408-5027 or 800-889-3358 (or email at tsupport@proquest.com or
http://support.proquest.com/ ) to find a solution.
Adding Supplements to the metadata in Excel and Moving them to the PDF Directory
Rename supplement files to a simple name that makes sense. Add their file names to the spreadsheet files with || separating file names. Then move them to the main PDF directory (even if they're not PDF).
In the filename column, enter the names of any extra files to be loaded in the appropriate line. Separate it from the existing file with ||. In the dc.description column for these, add a note indicating that there's a supplement and it's format, eg "Include 1 .jpeg3 supplement". Move the supplement from the supplement folder to the pdf folder after it's added to the metadata.
Author Details
check the log and complete author details. Add authors not in the Google doc to separate spreadsheet to prepare them to be added.
Adding Publication Forms to the metadata in Excel:
Log in to https://www.etdadmin.com/main/home and download the licenses fromthe academic term that was just completed.
For each line in the spreadsheet:
Find that author’s publication form either in the downloaded publication forms or by searching ETD Adminstrator to find it.
Open the publication form. If the publication form file doesn't contain a publication form, or is blank, hand the work as as limited access item.
Ensure that the publication form has the correct title. Remember if there's an embargo.
Do Save as...
Replace any blank spaces in the file name with an underscore. Then replace everything between the author's name and .pdf with Open. eg.:
Dutrow-Daryl_Open.pdf |
Copy the new publication form file name and paste it into the filename__bundle:LICENSE__permissions:-r'Anonymous' column.
For limited access items, copy and paste the filename of the thesis or dissertation into the filename__permissions:-r'ScholarWorksUMBCIP'__primary:true column.
For open access and embargoed items, copy and paste the filename of the thesis or dissertation into the filename__permissions:-r'Anonymous'__primary:true column.
If there's an embargo still in effect, save the the title, author, and the date the embargo expires in an embargoe list to add the embargoes after load.
Completing the Licenses and Deleting the filename Column
For all limited access items, change the value in the dcterms.Access rights column to this: "Access limited to the UMBC community. Item may possibly be obtained via Interlibrary Loan through a local library, pending author/copyright holder's permission."
Delete the file name column.
Check
Search for any blank spaces in license file names and fix them.
Save your Excel file final version.
Run the python program that compares the file names in the spreadsheet against the actual file names in the PDF file. Correct any errors.
Save your sheet2 (you must be on it) as a .csv file. It must be the plain .csv format, NOT the .csv UTF8 format--if you use UTF8 the load won’t work. While on the "save as" screen, change the character encoding to UTF8 by using the tools drop-down, selecting web options, then encoding, and UTF8.
Close the csv in Excel, and open it with Notepad++. Save as a UTF 8 file (select from menu Encoding>UTF-8 >> save) and close.
Run the SAF builder Using Ubuntu (documentation here: https://github.com/DSpace-Labs/SAFBuilder) :
Be sure the csv is closed in all programs.
Open the csv in NotePad++, convert to UTF8 and resave it. The PySAF program won’t run if you don’t do this step.
Download the PySAF zip file from above and unzip it.
Find the PySAF Updated.exe file and run it.
Select your csv file in the first box.
Select your pdf directory in the 2nd box.
Select the location were you want the SAF directory to go in the last box.
Do not check any boxes.
Click “Create Archive.”
Run the Date Format Checker Script
The Script is here:
Instructions for running it is here:
Correct any incorrectly formatted dates, and rerun the SAFBuilder and this script until all dates are correctly formatted.
Run the Collection Mapping Program
Installation instructions are here: User Documentation for Mapping Script
The program is here:
The program will add collection files to each subdirectoy in the SAF directory with the handles of the collections the work will be mapped to.
Run the Script:
Double-click the mapping_script.py file to execute the script.
Alternatively, run the script from the command line using the following command:
python Mapping_file_creation.py
Browse and Select Files:
Click on the "Browse" button next to "Folder Path" to select the folder containing metadata files. (SimpleArchiveFormat folder)
Click on the "Browse" button next to "Excel File" to select the Collections Excel file.
Execute Script:
After selecting the folder and Excel file, click on the "Run Script" button.
The script will process the metadata files, map the collections, and generate collections files for each item.
View Logs:
Logs for the script execution are stored in a file named mapping.log which is in the same folder where script file is stored.
This log file provides information about any errors encountered during the process.
If the log shows XML errors, it failed to run on the items with the errors. Correct them using Notepad, and then rerun the mapping program.
If the 2nd log shows errors, they are collections that didn’t match the collection spreadsheet. If they are collection that should have been added, manually add their handle to the collection file being careful not to add any extra characters or returns to the file.
Script Completion:
Once the script completes execution, a message box will appear indicating the successful creation of files or any errors encountered.
Run the Empty or Missing Collection File Script
The script is here:
Instructions for running it are here:
Add any missing collections or collections files before moving on
Zip the Final SimpleArchiveFormat directory with collection files and send to UMCP for Load
Open Ubuntu and zip the final SimpleArchiveFormat directory by:
navigating to the directory that the final SimpleArchiveFormat directory is inls
, using ls to view the contents of the directory you're in and cd to change directories.
Zip the SimpleArchiveFormat directory using the command zip -r myfilename.zip SimpleArchiveFormat/
Send the zipped SAF directory to MD-SOAR help, mdsoar-help@umd.edu, requesting that they load it.
