transcription-project.gmi
I will use this page to document my ongoing transcription project
Overall amount of footage: almost 8000 hours (7988hrs 12mins 2sec)
Total size of footage: 114terabytes
Footage was originally stored on 10 16terabyte drives
The Purpose Of This Project
The purpose of this project is to use intelligent tools to automatically transcribe the almost 8000 hours of rehearsals, meetings, and theatrical productions staged in the Tower Wall Community Theater. It was estimated these recordings were made from 2014 thru 2019 when the theater was closed.
Why Transcription
Because of the massive amount of footage, it will be helpful to first automatically or semiautomatically transcribe the footage into text to make searching for particular moments easier. The transcript will be accompanied by timecodes so that the archivist can easily retrieve clips as needed
Why did you take on this particular project?
I chose this project because in high school theater was important to me and I performed in several plays. I have also acted in a one-act here during freshman year. The Tower Wall theater is less than fifty miles from my hometown, though I never saw a play or musical there. This is my way of giving back to a form of art that I love.
Possible sorting methods
I intend to approach this project in a way that makes the resulting data easiest to sort and access. I wonder if I can create an automation that can match rehearsals together and collect them into one subcategory based on similarity of content. I am also interested to see if I can automatically sort out non-rehearsal/play material such as meetings based on textual content.
The camera
The camera was mounted to one side of the light booth. All footage I have viewed so far comes from this camera mounted at the same angle. I wonder if a future project could use differences in the mostly similar data to further sort footage, if that sort of sorting is even needed as this is not that big of a project. But maybe it is because theater is important. I don't know.
Identifying voices
This will be a tricky challenge. I want my transcription module to be able to identify distinct voices and ascribe them to specific performers. It was estimated that around 200 individuals participated in the productions at Tower Wall from 2014-2019. This means my module will have to recognize up to 200 different voices when fed the footage
I will also manually tag each recognized voice with the individual's name, which I will source from an archive of programs given to me.
My goal to start off is to find a play with less than five performers and use it as a testing ground for my module.
molly.flounder.online/