Is there even an easy, seamless vision assistant program?
I read textbooks on my PC and ideally i want a program with a normal chat environment where i can just hammer in questions about what's currently on my screen. Example: I'm working on a PDF, underline or circle things... and then just type "what does this sentence mean?", you get it.
I do NOT want to manually screenshot, navigate to the folder, drag the picture into the environment and then also have to type the question. It should also naturally be aware that the conversation is about what's on the screen, so i don't have to steer it with "make a screenshot; use your vision capabilites" etc.
I tried multiple different MCP in LM Studio, none of them were great...or worked :/
Someone said AnythingLLM has this function but i couldn't find it.
Is there a good solution?
Thank you in advance! :)