← Back to sungyongcho.com

A First Day of Voice-Directed AI Work

about 5 min read
#voice#ai-workflow#field-notes

I only began directing work on this site's logs page by voice yesterday evening, around 8 PM, after the microphone arrived. Every earlier record was typed. It is a very short change in practice, but it let me test a question I had been thinking about intensely since roughly last Thursday.

Why I bought a microphone

Typing has real advantages. It supports careful writing, gives thoughts room to settle, and remains especially valuable when prose needs direct editing. Pausing at a keyboard can itself be part of forming a thought.

But when an idea is developing quickly or I need to keep explaining direction at the pace of AI work, the act of input can sometimes become the bottleneck. I could already be at the next decision while I was still translating the previous one into a typed instruction. I bought the microphone because I wanted to get those freely developing thoughts out with less interruption.

I had seen ChatGPT Voice and had seen other developers use microphones in their work. That did not tell me whether it would suit my own workflow. I wanted to try it personally, not to follow a trend, but to see whether it made a practical difference to the work already in front of me.

Building this logs page by voice

Much of the current logs page was shaped through spoken direction: its layout, metadata, the timeline and navigation of the long PetV3 post, and the revisions that followed. I captured screenshots myself. Then I described by voice what should move, which area felt too wide, which information should sit nearer the title, and where the paths back to the list or between posts should remain visible.

That was where voice felt unexpectedly natural. Looking at a page and adding the next small judgment—make this quieter, move that upward, give this less weight—was easier than repeatedly converting each observation into a typed instruction. The loop still involved taking screenshots, explaining, and checking the result. Voice did not read the screen perfectly or produce a finished layout in one request. But it reduced one layer of translation while I was iterating on visual direction.

Writing and layout did not respond the same way

The clearest early difference was between page work and prose. Voice direction worked well for page configuration, layout, and visual choices that needed several quick adjustments. I could look at the result, state the next change, and keep the direction moving.

Prose quality still benefits from direct reading and editing. Voice can help get an initial thought out, but rhythm, word choice, paragraph order, and whether a sentence reaches too far are easier to judge with the text in front of me. Voice has not replaced writing; it has made the entrance into a thought wider.

That distinction is useful on its own. It suggests that one input method does not need to serve every part of the work. The speed and kind of attention required to direct a visual layout are different from the attention required to finish a paragraph.

Moving between more than one workstream

During this period I was not looking at one page or one thread alone. There were moments when I needed to move between more than one conversation or project flow, observe what each one was doing, and decide what direction was next. This was not perfect parallel work. It is still easy to lose context when attention shifts.

But when the direction was clear, switching attention and observing separate workstreams was less damaging than I expected. While one thread was waiting, I could inspect another screen or document and decide what to ask next. Returning to the first thread, I could often continue the thought verbally without rebuilding the whole instruction from scratch. This is not a measured productivity gain—only an early personal observation about coordinating more than one moving thread.

A first feeling for matching model to task

For this work I used Terra at low reasoning effort, and it produced good results for routine layout and document tasks. Moving page structure, aligning metadata, and revising document presentation did not always appear to need the most expensive model or the longest reasoning pass.

There may still be judgments that call for stronger, Fable-class models. But low-risk everyday work may not need to be handled at the same level. Matching a task to an appropriate model left me with a tentative sense that usage can be conserved without treating every request as a high-stakes reasoning problem. This is a personal rule I am still learning, not a proven best practice or a prescription for anyone else.

What this changed in my work

Typing still helps me shape a thought, but the bottleneck of getting a quickly developing thought out has eased. Prose still needs direct reading and editing, while voice felt especially natural for repeated page-setting and layout adjustments. It also helped me keep the next direction available while moving between separate workstreams.

Matching Terra at low reasoning effort to this kind of routine work looks, tentatively, like one way to use only the capacity a task needs. This is still a first record from roughly a day of voice-directed work. It is not a conclusion that voice is better, nor an answer for model allocation. It has simply made it easier to see what changes when thoughts can be expressed more quickly and the conversation with an AI can keep moving.