Was ist gwf-carbon-txt-crawler?
Der Webrobot gwf-carbon-txt-crawler indexiert und analysiert Inhalte von Webseiten. Er zeigt sich meistens mit der IP Adresse 46.62.134.217 und unter Verwendung des User Agent gwf-carbon-txt-crawler/1.0 (Web Crawler for the Green Web Foundation to measure the adoption of the carbon.txt initiative; https://www.thegreenwebfoundation.org/green-web-checker-faq/; support@greenweb.org). Mit 0.0004% Marktanteil ist gwf-carbon-txt-crawler auf Platz 237 der aktivsten Webrobots im Internet.
„carbon.txt is a single recognisable location on any web domain for public sustainability data relating to that company. The open source ecosystem of syntax, apis & plugins is free to use and build on. Be part of the movement for open, discoverable, comparable, and traceable sustainability data on the web. Watch our presentation of carbon.txt at State of the Open, London on YouTube. Listen to Chris Adams talking about carbon.txt on the Green.io podcast. Trusted by Organisations across the globe are using carbon.txt and changing the norms for disclosing sustainability data. What data can you share and find? AI model cards Model cards are overviews of how a particular AI model was designed and trained, and can contain structured information on CO2 equivalent emissions incurred during training. CSRD reports The Corporate Sustainability Reporting Directive (CRSD) is an EU law requiring certain companies to disclose data in accordance with the European Sustainability Reporting Standards (ESRS). Certificates Examples include Renewable Energy Certificates (RECs), Power Purchase Agreements (PPAs) or public contracts that show use of 100% green energy. TCS estimates Technology Carbon Standard (TCS) carbon emission estimates are based on categorised mappings of an organisation’s use of digital and be shared as structured JSON data. Our free and easy to use tools are there to help you get started. Get started Key features Surface data Bring data reported in organisational sustainability reports and disclosures into the public domain in a format that is easily read by both humans and machines. Extendable syntax A simple, extendable syntax that can be used to describe a wide range of sustainability data, from carbon emissions to social impact, and be easily updated as new data types and mandated reporting requirements emerge. Plugin architecture A plugin architecture that allows for the easy integration of carbon.txt validation tools into existing platforms and software, enabling discovery and validation of sustainability data at scale across a range of use cases. Get started Carbon.txt makes sustainability data easier to discover and use on the web. Carbon.txt is a single, discoverable location on any web domain for public, machine-readable, sustainability data relating to that company. It’s a web-first, connect not collect style approach, of most benefit to those interested in scraping the structured data companies have to publish according to national laws. Designed to be extended by default, we see carbon.txt becoming essential infrastructure for sustainability data services crunching available numbers and sharing the stories it can tell. Follow the guide Use the builder Use the validator Carbon.txt is a"
— Offizielle Beschreibung des Betreibers
Technische Einordnung von gwf-carbon-txt-crawler
gwf-carbon-txt-crawler wurde in Webserver-Logs als Bot oder Crawler erkannt. Die wichtigsten technischen Hinweise findest du auf dieser Seite: bekannte User-Agents, beobachtete IP-Adressen, Aktivitätsdaten und passende robots.txt-Regeln.
Für eine konkrete Entscheidung solltest du zusätzlich prüfen, welche URLs gwf-carbon-txt-crawler abruft, wie häufig die Zugriffe sind und ob der Bot deine robots.txt-Regeln respektiert.
Gefahreneinschätzung und Bewertung
Sollte man gwf-carbon-txt-crawler blockieren?
Prüfe zuerst Zugriffshäufigkeit, aufgerufene URLs und User-Agent. Danach kannst du entscheiden, ob eine Blockierung sinnvoll ist.
robots.txt – gwf-carbon-txt-crawler blockieren
Füge diese Zeilen in deine robots.txt ein, um gwf-carbon-txt-crawler den Zugriff auf deine Website zu verwehren:
User-agent: gwf-carbon-txt-crawler
Disallow: /
Du kannst den Zugriff auch gezielt einschränken, statt ihn komplett zu blockieren:
User-agent: gwf-carbon-txt-crawler
Disallow: /wp-admin/
Disallow: /wp-includes/
Allow: /
Häufige Fragen zu gwf-carbon-txt-crawler
Ist gwf-carbon-txt-crawler gut oder schlecht?
Das hängt vom Einsatzzweck ab. gwf-carbon-txt-crawler ist als Web-Crawler eingeordnet. Entscheidend sind Nutzen, Serverlast, Crawl-Verhalten und ob der Bot deine robots.txt-Regeln respektiert.
Wie erkenne ich gwf-carbon-txt-crawler in Server-Logs?
Suche nach dem User-Agent-Namen gwf-carbon-txt-crawler. Ein beobachteter User-Agent ist gwf-carbon-txt-crawler/1.0 (Web Crawler for the Green Web Foundation to measure the adoption of the carbon.txt initiative; https://www.thegreenwebfoundation.org/green-web-checker-faq/; support@greenweb.org). Vergleiche ausserdem IP-Adressen, Zugriffsmuster und aufgerufene URLs.
Reicht robots.txt zum Blockieren?
robots.txt ist ein Hinweis für regelkonforme Crawler. Unerwünschte oder aggressive Bots können diese Regeln ignorieren. In solchen Fällen helfen zusätzlich Firewall-Regeln, WAF-Regeln oder Blockierungen im Hosting/CDN.
Kann ein Bot seinen User-Agent fälschen?
Ja. Ein User-Agent ist leicht zu fälschen. Für wichtige Entscheidungen solltest du zusätzlich IP-Adresse, Reverse-DNS, Zugriffsmuster, Häufigkeit und aufgerufene URLs prüfen.
IP-Adressen 2 bekannte IPs
Diese IP-Adressen wurden bisher von gwf-carbon-txt-crawler verwendet:
46.62.134.217
77.42.93.194
User Agents
Mit diesen User-Agent-Strings identifiziert sich gwf-carbon-txt-crawler:
gwf-carbon-txt-crawler/1.0 (Web Crawler for the Green Web Foundation to measure the adoption of the carbon.txt initiative; https://www.thegreenwebfoundation.org/green-web-checker-faq/; support@greenweb.org)gwf-carbon-txt-crawler/1.0 (Web Crawler for the Green Software Foundation to measure the adoption of the carbon.txt initiative; https://carbontxt.org/; support@greenweb.org)